LLM Research

LLM Research

Our LLM research teams investigate the safety, clinical accuracy, and patient impact of healthcare AI models, so that artificial intelligence improves access to care as it becomes increasingly capable.

You are exceptional at software.
Now build the models behind it

LLM Research Papers Filter

TRADE: Transducer-Augmented Decoder for Speech LLM

Most speech AI models can’t respond while you’re still talking — they lose track of the audio as it arrives, so real-time conversation breaks down. TRADE pairs the language model with a streaming component that follows your words live, so one model can keep up in real time and know exactly when you’ve finished speaking....

Read Case Study

TRADE: Transducer-Augmented Decoder for Speech LLM

Most speech AI models can’t respond while you’re still talking — they lose track of the audio as it arrives,

Read More

A Simple Plug-in for Improving Eviction-Based KV Cache Compression

A simple inference add-on that lets AI models handle long conversations with far less memory.

Read More

Generative Artificial Intelligence-Driven Voice Assistance for Patient Education in Ophthalmology

Patients and eye doctors gave high marks to a voice AI that teaches patients about eye-injection treatment.

Read More

Crafting Reversible SFT Behaviors in Large Language Models

A way to install a new behavior in an AI model — and cleanly switch it off later.

Read More

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts

Makes the newest class of large AI models run significantly faster on the same hardware.

Read More

MixRAG: Mixture-of-Experts Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering

Helps AI answer complex questions by pulling knowledge from many angles at once instead of one.

Read More

ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection

Teaches AI models to keep the right memories — and drop the rest — during long conversations.

Read More

Perfecting Human–AI Interaction at Clinical Scale: Turning Production Signals into Safer, More Human Conversations

What 115 million real patient conversations taught us about making healthcare AI safer and more human.

Read More

HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue

The first benchmark measuring whether AI can genuinely emotionally support someone — scored side by side with humans.

Read More

ORION: Teaching Language Models to Reason Efficiently in the Language of Thought

Teaches AI to think in shorthand instead of essays — same answers, dramatically faster.

Read More