Expert judgment,
structured for
post-training.
AI labs need domain-expert graded data to train and evaluate models in medicine, law, and finance. We collect it, structure it, and deliver it in formats that go directly into post-training pipelines.
The problem
Crowdsourced feedback breaks down in specialized domains.
Post-training (RLHF, DPO, supervised fine-tuning, RL) shapes how a model behaves on a specific task. The quality of that process depends entirely on the quality of human feedback driving it.
For general tasks this is straightforward. Crowdworkers can judge whether a summary is readable. But in medicine, law, or finance, the gap between a plausible answer and a dangerous one isn’t visible to a generalist. A first-year resident can miss what a cardiologist catches in seconds. Most AI labs either skip domain-specific post-training, or run it with evaluators who aren’t qualified to distinguish good outputs from subtly wrong ones.
The result is models that sound authoritative but fail where it matters most.
What we build
Expert preference datasets
We source credentialed domain experts (physicians, attorneys, scientists, financial analysts) and run structured annotation workflows for RLHF and DPO. A doctor comparing two diagnostic plans. A lawyer ranking two contract analyses. Each judgment is logged with structured reasoning and delivered in standard formats.
Domain evaluation frameworks
We build evals that go beyond multiple-choice benchmarks. Whether a clinical recommendation is appropriate requires someone who has actually made clinical decisions. We design scenario-based tasks, expert-graded rubrics, and scoring pipelines for domains where "correct" can't be reduced to an answer key.
RL environments and reward infrastructure
For agentic tasks like a clinical documentation system or a legal research tool, you need a reward function. There's no unit test for "is this clinical reasoning sound." We build the environments and expert-in-the-loop reward pipelines that let labs run policy gradient training on complex professional tasks.
Research
WHBench: A Women’s Health Benchmark
for Evaluating Frontier LLMs
Paper ↗Across 3,100 scored responses, no model mean exceeds 75%, with substantial safety and equity gaps across clinical topics. No widely adopted benchmark had previously evaluated frontier LLMs on women’s health.
Modalities