Research · Sep 2026
Schedule of Activities Extractor
A provenance-aware document pipeline for locating, extracting, stitching, and validating clinical-trial Schedule of Activities tables.
The problem
Clinical-trial protocols hide wide, multipage Schedule of Activities tables inside long PDFs. Flattening them to text loses hierarchy, footnote links, ambiguous cells, and page provenance—the exact details a downstream clinical workflow needs to trust.
What I built
I built a Python pipeline and upload UI that locates candidate pages, renders them for multimodal extraction, deterministically stitches multipage tables, and emits a graph schema that preserves headers, verbatim cell values, footnotes, provenance, and unresolved ambiguity.
Location and extraction are separate problems
A model cannot extract a table it never sees. The first stage finds likely Schedule of Activities pages across the entire protocol, and the extraction stage runs only on those images. Keeping the stages separate made it possible to measure page recall independently and reach 12/12 on the reference set and 8/8 on unseen holdouts.
Deterministic stitching protects the source
The table may repeat headers, continue rows across pages, or move footnotes to the final fragment. The stitcher joins extracted fragments with deterministic rules instead of prompting a model to summarize them, which preserves verbatim cell values and keeps every merge inspectable.
Ambiguity is data
The graph schema records unresolved structures and ambiguous cells rather than silently forcing them into a convenient shape. It also retains hierarchical headers and footnote targets. That makes the output suitable for review-heavy clinical workflows where an explicit unknown is safer than a clean but invented answer.
Outcome
- Located 12/12 reference pages across five protocols and 8/8 pages across three unseen holdout protocols
- Recovered 31/31 and 37/37 assessments on two hand-verified protocols
- Preserves hierarchical headers, verbatim cells, footnote text and linkage, ambiguity, and provenance
- Compared text-layer, Gemini, and OpenAI extraction paths and maintains 28 no-API regression tests
Stack
Data flow
Long clinical protocol to provenance-bearing activity graph
- 1
Protocol upload
FastAPI/Uvicorn service accepts full clinical-trial PDFs
- 2
Page locator
Scores and selects candidate Schedule of Activities pages
- 3
Page renderer
Produces page images while retaining page-level source identity
- 4
Multimodal extraction
Reads rows, columns, cells, hierarchy, and footnote markers from page images
- 5
Deterministic stitcher
Joins multipage table fragments without asking a model to rewrite source values
- 6
Provenance graph
Links every value, header, footnote, ambiguity, and unresolved structure back to evidence
- 7
Evaluation
Reference, holdout, hand-verified assessment counts, and 28 offline regression tests
This diagram is generated from portfolio.ts. Edit the `architecture` field to change it.
Keep reading
Peptide AI Product, Mobile, and RAG Platform
Cross-platform product engineering, lifecycle automation, retrieval, safety evaluation, and a privacy-scoped semantic cache.
Case studySynthDrive
You describe a driving scenario in plain English, and it generates a 3D world, runs object detection on it, and scores how hard that scene is to perceive.
Case studyEgoSocial (Meta Project Aria research)
A UMD research proposal for an egocentric dataset that pairs Aria Gen 2 sensor streams with validated psychological ground truth, plus the world model and behavioral model trained on it.
Case studyWant the deeper version of this?
I am happy to walk through the tradeoffs, the failure modes, and what I would do differently.