The research team found that the performance of the AI model improved after compression.
Coinpaper
10h ago
Ai Focus
According to Multiverse Computing, their new method can compress GPT-OSS from 120B to 60B, with a 4-bit reduction, and in most tests, it outperforms the corresponding 60B full-precision version.
Helpful
No.Help

一家研究团队提出了一种新的 AI 模型压缩方法,试图打破“模型越小、性能越弱”的常见取舍。Multiverse Computing 于 8 月 25 日在 Hugging Face 博客介绍称,他们把 OpenAI 开源的 GPT-OSS 120B 压缩到 600 亿参数,并进一步压到 4-bit 表示后,模型在多项测试中的表现反而优于对应的 60B 全精度版本。

9 项测试中赢下 7 项

研究团队将这套方法命名为 Quantization-Aware Healing。按其披露,压缩后的 60B 模型在 9 项基准测试中有 7 项超过“未量化”的 60B 版本,也就是通常被视为更高质量的对应模型。

这并不意味着压缩版已经全面超过原始的 120B 大模型。文章提到,120B 原版在多数对比中仍然更强,但压缩后的版本跨过了一个此前不常见的门槛:在更低内存占用和更少参数下,仍能取得更好的局部测试成绩。

  • 原始模型:GPT-OSS 120B
  • 压缩后规模:60B 参数
  • 内存表示:4-bit

关键在于学习对象变了

常见做法是,先把大模型缩小,再让更小的模型去模仿这个“中间版本”。问题在于,这个中间版本本身已经在压缩过程中损失了一部分能力。这样训练出来的小模型,通常只能逼近这份“打折后的答案”。

这次的方法改了训练目标。研究团队没有让小模型继续对齐缩水后的中间版本,而是直接让它向原始、未压缩的大模型学习输出结果。换句话说,压缩步骤不再只是性能损耗,而被当作一次额外的监督训练机会。

对本地部署更有吸引力

如果这一结果能在更多模型上复现,意义主要在部署端。更小的参数规模和更低的位宽,意味着模型运行时需要的显存和电力都会下降。文章称,这个“修复后”的版本大约只需要原模型四分之一的内存,同时参数量减半。

这会直接影响模型能否从数据中心走向桌面设备,甚至进一步进入移动端。对预算有限的实验室、独立开发者和本地部署用户来说,推理成本下降本身就是重要变化。

已发布开源权重

Multiverse Computing 表示,压缩后的 Hypernova-60B 已以开放权重形式发布在 Hugging Face,用户可以下载运行。不过,生成这一模型所用的压缩工具仍属专有工具,完整流程目前并未完全开源。

补充信息:这项测试目前只覆盖 GPT-OSS,并未扩展到 Llama、Qwen 或 Mistral 等其他主流模型家族,方法的通用性仍待后续验证。

Tip
$0
Like
1
Save
0
Views 54
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Ventures Platform Expands the Layout of its Second Phase Fund in Africa
Ventures Platform Launches a second African fund to expand regional coverage and focus on AI's cost restructuring capabilities in the African market.
TechCrunch
·2026-08-26 15:13:56
12
web3: Foreign media: Bitcoin falls back after breaking above $80,000, with $82,830 becoming a key level
Bitcoin surges briefly before pulling back; foreign media cites analysts saying $82,830 is a key level for judging trend changes.
CoinPedia
·2026-08-26 10:03:49
45
Jetson Orin Nano 2 Entry-Level Robots: The Conditions for Implementation Behind 78 TOPS and 15 Watt Energy Efficiency
NVIDIA released on August 25th the Jetson Orin Nano 2 robot computer, aimed at entry-level edge AI, robots, delivery and inspection drones, as well as visual AI systems. The official specifications include 78 trillion operations per second, 8GB of memory, and an 8-core Arm CPU; compared to Jetson Orin Nano Super, the inference performance has been doubled, while the physical dimensions remain unchanged. In 15-watt mode, it operates with 40% less power consumption while maintaining the same performance.
CoinMeta
·2026-08-26 09:56:15
34
OpenAI Announces Its First Self-Developed Inference Chip Jalape: There Are Actual Test Results, But This Does Not Mean It Can Completely Replace GPU
On August 25th, OpenAI disclosed for the first time the measurement results of its self-developed inference chip Jalape. The company stated that, using a public InferenceX benchmark with a capacity of 120B GPT-OSS, this chip achieved a higher peak throughput per kilowatt and lower token latency compared to the commercial systems involved in the comparison; it also performed strongly on DeepSeek R1 and Kimi K2. OpenAI indicated that they currently possess a working first-party chip, and subsequent generations are also under development.
CoinMeta
·2026-08-26 09:55:52
26
Robot AI and startup Generalist valued at $3 billion
Generalist reportedly completed nearly $200 million in additional financing, raising its valuation to $3 billion, reflecting continued capital betting on the AI model of general-purpose robots.
TechCrunch
·2026-08-26 08:48:54
34
View More