PaRLA: base Llama 70B vs PaRLA

Real TCGA pathology reports abstracted by the base model and by PaRLA, with the GPT-5.5 (Codex) judge's verdict. Precomputed from committed records. Model: huggingface.co/AliKhajegiliM/PaRLA

Base Llama 70B

PaRLA

Cases are drawn from the released judgments.jsonl (500 TCGA reports). "Facts missed" are the judge's recorded omissions for each output. Reproduce the aggregate statistics with analyze_judgments.py.