Skip to content
All projects

Research · Sep 2026

Schedule of Activities Extractor

A provenance-aware document pipeline for locating, extracting, stitching, and validating clinical-trial Schedule of Activities tables.

The problem

Clinical-trial protocols hide wide, multipage Schedule of Activities tables inside long PDFs. Flattening them to text loses hierarchy, footnote links, ambiguous cells, and page provenance—the exact details a downstream clinical workflow needs to trust.

What I built

I built a Python pipeline and upload UI that locates candidate pages, renders them for multimodal extraction, deterministically stitches multipage tables, and emits a graph schema that preserves headers, verbatim cell values, footnotes, provenance, and unresolved ambiguity.

Location and extraction are separate problems

A model cannot extract a table it never sees. The first stage finds likely Schedule of Activities pages across the entire protocol, and the extraction stage runs only on those images. Keeping the stages separate made it possible to measure page recall independently and reach 12/12 on the reference set and 8/8 on unseen holdouts.

Deterministic stitching protects the source

The table may repeat headers, continue rows across pages, or move footnotes to the final fragment. The stitcher joins extracted fragments with deterministic rules instead of prompting a model to summarize them, which preserves verbatim cell values and keeps every merge inspectable.

Ambiguity is data

The graph schema records unresolved structures and ambiguous cells rather than silently forcing them into a convenient shape. It also retains hierarchical headers and footnote targets. That makes the output suitable for review-heavy clinical workflows where an explicit unknown is safer than a clean but invented answer.

Outcome

  • Located 12/12 reference pages across five protocols and 8/8 pages across three unseen holdout protocols
  • Recovered 31/31 and 37/37 assessments on two hand-verified protocols
  • Preserves hierarchical headers, verbatim cells, footnote text and linkage, ambiguity, and provenance
  • Compared text-layer, Gemini, and OpenAI extraction paths and maintains 28 no-API regression tests

Stack

PythonFastAPIPyMuPDFpdfplumberMultimodal LLMsGraph schemasRegression testing

Data flow

Long clinical protocol to provenance-bearing activity graph

  1. 1

    Protocol upload

    FastAPI/Uvicorn service accepts full clinical-trial PDFs

  2. 2

    Page locator

    Scores and selects candidate Schedule of Activities pages

  3. 3

    Page renderer

    Produces page images while retaining page-level source identity

  4. 4

    Multimodal extraction

    Reads rows, columns, cells, hierarchy, and footnote markers from page images

  5. 5

    Deterministic stitcher

    Joins multipage table fragments without asking a model to rewrite source values

  6. 6

    Provenance graph

    Links every value, header, footnote, ambiguity, and unresolved structure back to evidence

  7. 7

    Evaluation

    Reference, holdout, hand-verified assessment counts, and 28 offline regression tests

PythonFastAPIPyMuPDFpdfplumberMultimodal LLMsGraph schemasRegression testing

This diagram is generated from portfolio.ts. Edit the `architecture` field to change it.

Want the deeper version of this?

I am happy to walk through the tradeoffs, the failure modes, and what I would do differently.

Email me