Backed by Y Combinator

Expert judgment,
structured for
post-training.

AI labs need domain-expert graded data to train and evaluate models in medicine, law, and finance. We collect it, structure it, and deliver it in formats that go directly into post-training pipelines.

Scroll
Healthcare·Life Sciences·Finance·Legal·Pharmaceuticals·Clinical Research·Drug Discovery·Radiology·Healthcare·Life Sciences·Finance·Legal·Pharmaceuticals·Clinical Research·Drug Discovery·Radiology·Healthcare·Life Sciences·Finance·Legal·Pharmaceuticals·Clinical Research·Drug Discovery·Radiology·Healthcare·Life Sciences·Finance·Legal·Pharmaceuticals·Clinical Research·Drug Discovery·Radiology·

The problem

Crowdsourced feedback breaks down in specialized domains.

Post-training (RLHF, DPO, supervised fine-tuning, RL) shapes how a model behaves on a specific task. The quality of that process depends entirely on the quality of human feedback driving it.

For general tasks this is straightforward. Crowdworkers can judge whether a summary is readable. But in medicine, law, or finance, the gap between a plausible answer and a dangerous one isn’t visible to a generalist. A first-year resident can miss what a cardiologist catches in seconds. Most AI labs either skip domain-specific post-training, or run it with evaluators who aren’t qualified to distinguish good outputs from subtly wrong ones.

The result is models that sound authoritative but fail where it matters most.

What we build

01

Expert preference datasets

We source credentialed domain experts (physicians, attorneys, scientists, financial analysts) and run structured annotation workflows for RLHF and DPO. A doctor comparing two diagnostic plans. A lawyer ranking two contract analyses. Each judgment is logged with structured reasoning and delivered in standard formats.

02

Domain evaluation frameworks

We build evals that go beyond multiple-choice benchmarks. Whether a clinical recommendation is appropriate requires someone who has actually made clinical decisions. We design scenario-based tasks, expert-graded rubrics, and scoring pipelines for domains where "correct" can't be reduced to an answer key.

03

RL environments and reward infrastructure

For agentic tasks like a clinical documentation system or a legal research tool, you need a reward function. There's no unit test for "is this clinical reasoning sound." We build the environments and expert-in-the-loop reward pipelines that let labs run policy gradient training on complex professional tasks.

Research

WHBench: A Women’s Health Benchmark for Evaluating Frontier LLMs

Paper ↗
47
Clinical scenarios
22
Models evaluated
23
Rubric criteria
3,100
Scored responses

Across 3,100 scored responses, no model mean exceeds 75%, with substantial safety and equity gaps across clinical topics. No widely adopted benchmark had previously evaluated frontier LLMs on women’s health.

Modalities

Expert knowledge comes in many forms.

TextStructured reasoning, preference pairs, rubric evaluations.
VoiceSpoken expert reasoning during clinical and legal encounters.
VideoEgocentric recordings of procedural tasks: surgery, lab work, physical examination.
MultimodalCross-modal datasets combining text, image, audio, and action.