OneTriangle
@onetriangle.ai
OneTriangle optimizes multi-model inference by transferring KV cache between models, cutting redundant prefill cost.
OneTriangle's Company Logos
OneTriangle's Brand Colors
Hex Code
Color name
RGB
HSL
CMYK
#0B1516
Aztec
11, 21, 22
185, 33, 6
50, 5, 0, 91
#FFFFFF
White
255, 255, 255
0, 0, 100
0, 0, 0, 0
About OneTriangle
OneTriangle is an inference technology company focused on making large language models faster and less costly to run. Its inference engine transfers key-value (KV) cache information between models, allowing a smaller model to process an input before a larger model generates the response. This approach is designed to reduce redundant prefill work and lower costs, particularly for long-context requests.
OneTriangle reports up to 20% lower inference costs and time overall; for a 100,000-token context using Llama 3. 1 8B to 70B, it cites a 6. 3-times faster first token and 84% lower prefill cost.
The company validates model pairings against quality and latency requirements, using standard inference when a pairing does not meet its criteria. OneTriangle also provides managed serving for open-weight models, handling GPU infrastructure and setup so teams can integrate the service with a configuration change. Supported model families listed on its site include Llama, Qwen, Mistral, Gemma, and DeepSeek.
The company shares research on cross-model KV cache transfer and offers product demonstrations
Brand industry
Computers Electronics and Technology
Company type
Suggest company type
Year founded
Suggest founded year
Company size
Suggest company size
