Request a demo
Evaluate Relay locally with a small-model setup.
See UnimAI in action
A short walkthrough of Relay moving KV cache state across a model boundary — the cross-model transfer that replaces a full prefill on handoff.
Demo overview
You only prefill once. We translate. Your model takes it from there.
In this demo, UnimAI Cross-Model KV Cache Transfer converts the KV cache generated by Qwen2.5-1.5B-Instruct into a representation that Gemma-2-2B-it can immediately consume and continue decoding from.
Instead of forcing Gemma-2-2B-it to repeat a full prefill, UnimAI transfers reusable computation already completed by Qwen2.5-1.5B-Instruct, reducing redundant prefill work and handoff latency.
Apply to try the demo and:
- Enter a prompt to chat or upload a TXT document for Q&A.
- Compare native Qwen2.5-1.5B-Instruct, native Gemma-2-2B-it, and KV-transferred Gemma.
- Inspect output quality and Prefill / Consume time side by side.
- Test long-context and reasoning workloads.
Stop recomputing what has already been computed. Transfer the intelligence forward.
Tell us a little about you and what you'd like to evaluate. Access is reviewed by hand.
