Prime Intellect Inference Platform Launched: Nearly 1 trillion Token Calculations Per Day Internally
2026-10-04 23:31:03
According to CoinMeta, the Prime Intellect inference platform has been officially launched, processing nearly 1 trillion tokens per day internally. The platform has released an open-source model inference service called prime inference, allowing users to call models as needed and supporting long-term stable operation with compute power locking. The interface is compatible with OpenAI API, and the first publicly deployed model is glm-5.3. Since January of this year, Prime Intellect has been running large-scale deployments for customers, with a focus on optimizing the long-context load of agent. During testing, the latency between p90 and token was reduced by nearly 40%. At the same time, Prime Intellect has compressed the kv cache, increasing the number of token that each decoder can cache from 1.09 million to 1.63 million under the same amount of video memory, which is a boost of about 50%.
Bullish 1
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.