Why It Is Harder Than It Looks
Scientific PDFs are page-layout documents optimized for print, not for data reuse. Figures are embedded as images — rasterized or vector — with no machine-readable label identifying them as figures. Tables may be rendered as formatted text, as images, or as a mix of both. There is no universal standard across publishers.
The difficulty compounds with multi-panel figures (panels A, B, C, D arranged in a single figure block), supplementary appendices published as separate files, and scanned pages in older articles. A reader can instantly recognize a Kaplan-Meier curve; automated systems historically struggled to locate it reliably within the page.
What You Actually Need to Extract
- Figures — Kaplan-Meier survival curves, forest plots, waterfall plots, flow cytometry, microscopy, pathway diagrams
- Tables — patient characteristics (Table 1), efficacy outcomes, safety data, response rates
- Data values from figures — reading OS rates, HR, or response percentages directly off a graph axis
- Supplementary figures — often the most detailed data, buried in appendix PDFs
For literature reviews and slide decks, the bottleneck is rarely finding the papers — it is getting the key visual evidence out of them quickly.
How AI-Powered Extraction Works
Modern AI vision models process each PDF page rendered as a high-resolution image. The model identifies the position of every figure on the page — distinguishing data figures from CONSORT flow diagrams, supplementary tables from main-text tables, and individual panels within a multi-panel figure block.
Once bounding boxes are detected, the corresponding region is cropped from the original PDF at full resolution. The result is a clean, publication-quality image file ready for a slide deck or a review document.
Beyond extraction, AI can also read the figures: identifying what type of figure it is, summarizing the result it shows, and flagging which figures are most relevant to a given research question.
Figures in Presentations: Copyright
Most publishers allow figure reuse for educational and non-commercial purposes with attribution. Open-access articles published under a CC BY licence explicitly permit reuse in any context. Always cite the original article and journal when reproducing a figure in a presentation or review.
Multi-Panel Figures: A Common Error
Many oncology figures arrange four to six panels in a single figure block. Extracting panels individually — without preserving the relationship between them — loses the narrative the authors built. The convention is to capture the complete figure as one unit, then annotate individual panels in the slide or document where you use it.
Extract figures from any medical PDF — automatically
Upload your papers to PubRank. AI scans every page, identifies all figures and tables, and extracts them at full resolution in one click.
Try PubRank free →