Doctors spend an agonizing amount of time reading through lengthy EMR notes, discharge summaries, and lab findings. Fatigue is real, and it delays critical decisions. The obvious fix is to pass the text to an LLM like ChatGPT and ask for a summary.
In a clinical setting, that fix is actively dangerous.
Generative LLMs paraphrase. If an AI hallucinates a symptom, invents a timeline, or subtly alters a drug dosage in the paraphrase, it could cost a life. You cannot ship a generative summarizer into clinical workflows and sleep at night.
The medical literature is full of cases where a single misplaced decimal or a swapped drug name caused an adverse event. Abstractive summarization — where the model generates new text — is the wrong tool for the job.
Extractive, not generative
The engine acts as a highly intelligent highlighter, not a writer. Instead of writing new text, it scores every sentence in the source document by clinical salience and returns only the most critical ones — verbatim from the source. No new tokens are generated. What the doctor reads is exactly what was written in the source, just filtered to what matters most.
Fine-tuned SciBERT — pre-trained on biomedical literature with its own SciVocab — powers the scoring. Deployed as a FastAPI inference API on Hugging Face Spaces.
Phase 1: the lazy baseline
I built an unsupervised K-Means pipeline using raw SciBERT embeddings to cluster sentences, picking the centroid of each cluster as a summary sentence. ROUGE-1 came in at 0.2631. Respectable for a baseline, useless for production.
The problem was fundamental: K-Means groups by mathematical distance in embedding space, not clinical importance. A sentence about the patient's parking arrangements could cluster with a sentence about a diagnosis just because their embeddings happened to be near each other. Context-blind clustering can't be trusted in a domain where context is the whole point.
Phase 2: the labels didn't exist, so I built them
This is where most people give up.
The ccdv/pubmed-summarization dataset on HuggingFace only had human-written abstracts — abstractive summaries, not extractive labels. I needed sentence-level 0/1 labels to train a supervised classifier, and none existed.
So I wrote a Greedy Matching algorithm. For every training record, it iterated through every source sentence and calculated how much the ROUGE score would increase if that sentence were included in a candidate summary. It hunted for the exact combination of original source sentences that best reconstructed the human-written abstract, and labeled each sentence 1 (include) or 0 (ignore). I ran it across 200 training records and now had a real supervised dataset.
Labeling strategy matters more than model architecture. Without the greedy matching step, there was no training signal, and no amount of fine-tuning would have moved the score.
Phase 3: BERTsum and the head swap
Standard SciBERT produces one embedding for a whole document — useless for sentence-level scoring. I implemented BERTsum architecture: inserted [CLS] and [SEP] tokens around every individual sentence, added interval segment embeddings so the model could distinguish adjacent sentences, then chopped off SciBERT's default masked-language-model head and attached a custom PyTorch nn.Linear classification layer I called BioExtractor.
Fine-tuned the whole thing on a Tesla T4 GPU in Google Colab. ROUGE-1 jumped from 0.2631 to 0.3782 — a 43.75% improvement over baseline. The model had learned to think like a clinician: prioritizing diagnoses, outcomes, and dosages over general fluff.
Phase 4: the deployment nobody warns you about
A model in a Colab notebook is useless in production. I built a FastAPI async inference API — but deploying a custom PyTorch architecture isn't like calling AutoModel.from_pretrained(). Standard HuggingFace pipelines don't know about my custom BioExtractor head.
I had to save the fine-tuned weights as a .pt file, upload it directly into the Hugging Face Space, and write a lazy-loading function that injects the weights into the SciBERT skeleton the moment the server boots. If I got the loading order wrong the weights wouldn't map to the layers correctly and the model would run but produce garbage. This was the debugging cycle I dreaded most — silent failures where the model 'worked' but was nonsense.
The one-line safety feature
One final touch, and this is the part I'm proudest of. Before returning the top N sentences, I sort them back into their original chronological index.
In medicine, the timeline is everything. A summary that puts the diagnosis before the symptom, or the treatment before the diagnosis, creates dangerous clinical confusion. A single sort() call turned out to be the most clinically important line of code in the whole project.
The most important engineering decisions aren't always the hardest. Sometimes safety is one line.
What I learned
Generative AI is not the right tool for every problem, and knowing when not to use it is more valuable than knowing how to prompt it well. The extractive-vs-generative decision was the whole project — the model architecture followed from it.
Labeling strategy beats model architecture. If your data isn't labeled for your actual task, you don't have a training problem, you have a labeling problem.
Deploying custom PyTorch architectures requires explicitly defining the model class skeleton in your serving code. The moment you deviate from vanilla HuggingFace, you're responsible for the loading pipeline.
What I owned
- Phase 1 — Built an unsupervised K-Means baseline using raw SciBERT 768-dimensional embeddings, establishing a ROUGE-1 score of 0.2631 as the starting benchmark and demonstrating why context-blind clustering isn't good enough for clinical use
- Phase 2 — Wrote a Greedy Matching algorithm to generate extractive sentence labels from the ccdv/pubmed-summarization dataset, which only contained abstractive abstracts — solving the labeling problem that would have blocked supervised training entirely
- Phase 3 — Implemented BERTsum architecture: inserted [CLS]/[SEP] tokens around every sentence, added interval segment embeddings, and attached a custom PyTorch classification head (BioExtractor) for binary salience scoring
- Fine-tuned the full model on a Tesla T4 GPU in Google Colab, achieving a ROUGE-1 of 0.3782 — a 43.75% improvement over baseline
- Phase 4 — Built a FastAPI async inference API with lazy-loading of custom .pt weights into the SciBERT skeleton, deployed to Hugging Face Spaces as a live inference endpoint
- Implemented chronological re-sorting of extracted sentences before returning the final summary — preserving clinical timeline integrity, which is safety-critical for medical use