Request a demo

Evaluate Relay locally with a small-model setup.

See UnimAI in action

A short walkthrough of Relay moving KV cache state across a model boundary — the cross-model transfer that replaces a full prefill on handoff.

Demo overview

You only prefill once. We translate. Your model takes it from there.

In this demo, UnimAI Cross-Model KV Cache Transfer converts the KV cache generated by Qwen2.5-1.5B-Instruct into a representation that Gemma-2-2B-it can immediately consume and continue decoding from.

Instead of forcing Gemma-2-2B-it to repeat a full prefill, UnimAI transfers reusable computation already completed by Qwen2.5-1.5B-Instruct, reducing redundant prefill work and handoff latency.

Apply to try the demo and:

  1. Enter a prompt to chat or upload a TXT document for Q&A.
  2. Compare native Qwen2.5-1.5B-Instruct, native Gemma-2-2B-it, and KV-transferred Gemma.
  3. Inspect output quality and Prefill / Consume time side by side.
  4. Test long-context and reasoning workloads.

Stop recomputing what has already been computed. Transfer the intelligence forward.

Tell us a little about you and what you'd like to evaluate. Access is reviewed by hand.

I'm reaching out as