Cohere Launches 2.4B Visual Mini-Model: Document Understanding Outperforms Ministral's 33B Model
2026-08-13 15:55:53
According to CoinMeta, Cohere has open-sourced its smallest visual language model to date, north, with 2.4B parameters and licensed under the Apache 2.0 license. This model is specialized in processing documents, tables, charts, screenshots, and OCR. It can handle images in their original proportions and resolutions without the need for compression first. In official tests, docvqa achieved a score of 92.1%, surpassing the 89.6% of Ministral (with 33B parameters) and the 73.2% of Gemma and E2B. Its visual positioning capability scored 73.2%, which is also higher than that of Ministral and Gemma. Although it performs excellently in document understanding and visual positioning, it is not the strongest model in terms of size; Qwen3.5-2B is still stronger in most general visual tasks, OCR, and multimodal benchmarks.
Source:Internet
This content is for market information only and does not constitute investment advice.
Follow WalletJYS official accounts to stay updated

Hot Articles
Refresh

'No longer a distant place': F2Pool Co-founder Chun Wang joins SpaceX's 2-year mission to Mars
05-22 18:25

Polymarket Targets Japan Approval Despite Gambling Laws
05-22 18:00

ZachXBT flags suspected exploit involving Polymarket's UMA adapter contract on Polygon
05-22 17:57

ZachXBT flags $520K Polymarket exploit on Polygon, team says funds are safe
05-22 17:24

Verus bridge exploiter returns 4,052 ETH, retains $2.8 million bounty: onchain analyst
05-22 17:24



