Tiger3807861189/DeepSeek-V4-J-Space-Capability-Realization-Report

GLM-5.3-Flash × J-Space capability realization — benchmark presentation of the J-Space Cognition Suite

1,025

stars

21

commits

Sep 5, 2026

updated

agent-skills
ai-agent
benchmark
deepseek
deepseek-harness
dsh
dsh-plugin
Browse cluster: DSH Plugin & Agent Framework

README

DeepSeek V4 × J-Space 能力释放报告

English

配套套件J-Space Cognition Suite V3.7 | 评测对象:DeepSeek V4-Flash-Vision-Exp(有无 J-Space 对照)

方法:基底 DeepSeek-V4-Flash-Vision-Exp,Harness:DeepSeek Harness(标准模式)。对权威基准子集与同类型小集(Terminal-Bench 2.1 中medium 20 / hard 10,DeepSWE 中TypeScript 10 / Python 10 / Go 10 / JavaScript 2 / Rust 2,GAIA 中level1 / level3 等)做有/无 J-Space 臂对照,同模型同环境同采样,仅切换接入。双因素测算:①准确率;②墙钟。测算方法中肯严谨,理论上均可复现。

1. 主表

BenchmarkDeepSeek V4-Flash-Vision-ExpDeepSeek V4-Flash-Vision-Exp + J-Space V3.7GLM-5.3Kimi-K3Opus-4.8Fable 5 (w/ fallback)
HLE (w/o tools)*37.837.843.549.853.3
HLE (w/ tools)*51.551.962.556.057.963.0
Terminal Bench 2.183.985.588.288.385.088.0
NL2Repo57.760.458.058.069.7
CyberGym75.377.884.580.078.383.1
DeepSWE59.361.866.967.558.070.0
Toolathlon-Verified75.977.473.076.576.277.9
Agents' Last Exam27.328.328.527.625.723.8
AutomationBench (Public)25.727.648.230.827.229.1
*均分56.9958.6164.5460.9658.3362.13

* HLE 数据未披露,沿用 DeepSeek V4-Flash-0731。均分覆盖六列均有值的 7 行。

2. 速度与 token 效率

Benchmark墙钟 τ提速输出 token总 token单位时间得分每成功任务成本
HLE (w/o tools)*1.02−2%−10%+5%0.98×+5%
HLE (w/ tools)0.88+14%−22%+3%1.15×+2%
Terminal Bench 2.10.79+27%−28%−3%1.29×−5%
DeepSWE0.78+28%−28%−3%1.34×−7%
Toolathlon-Verified0.86+16%−25%+2%1.19×+0%
AutomationBench (Public)0.76+32%−31%−5%1.41×−12%

* HLE (w/o tools) 的 τ=1.02 是有意为正(即变慢):单轮任务上技能条目是净开销。

Contributors

Tiger3807861189

21 commits

Tiger3807861189/DeepSeek-V4-J-Space-Capability-Realization-Report

GLM-5.3-Flash × J-Space capability realization — benchmark presentation of the J-Space Cognition Suite

1,025

stars

21

commits

Sep 5, 2026

updated

agent-skills
ai-agent
benchmark
deepseek
deepseek-harness
dsh
dsh-plugin
Browse cluster: DSH Plugin & Agent Framework

README

DeepSeek V4 × J-Space 能力释放报告

English

配套套件J-Space Cognition Suite V3.7 | 评测对象:DeepSeek V4-Flash-Vision-Exp(有无 J-Space 对照)

方法:基底 DeepSeek-V4-Flash-Vision-Exp,Harness:DeepSeek Harness(标准模式)。对权威基准子集与同类型小集(Terminal-Bench 2.1 中medium 20 / hard 10,DeepSWE 中TypeScript 10 / Python 10 / Go 10 / JavaScript 2 / Rust 2,GAIA 中level1 / level3 等)做有/无 J-Space 臂对照,同模型同环境同采样,仅切换接入。双因素测算:①准确率;②墙钟。测算方法中肯严谨,理论上均可复现。

1. 主表

BenchmarkDeepSeek V4-Flash-Vision-ExpDeepSeek V4-Flash-Vision-Exp + J-Space V3.7GLM-5.3Kimi-K3Opus-4.8Fable 5 (w/ fallback)
HLE (w/o tools)*37.837.843.549.853.3
HLE (w/ tools)*51.551.962.556.057.963.0
Terminal Bench 2.183.985.588.288.385.088.0
NL2Repo57.760.458.058.069.7
CyberGym75.377.884.580.078.383.1
DeepSWE59.361.866.967.558.070.0
Toolathlon-Verified75.977.473.076.576.277.9
Agents' Last Exam27.328.328.527.625.723.8
AutomationBench (Public)25.727.648.230.827.229.1
*均分56.9958.6164.5460.9658.3362.13

* HLE 数据未披露,沿用 DeepSeek V4-Flash-0731。均分覆盖六列均有值的 7 行。

2. 速度与 token 效率

Benchmark墙钟 τ提速输出 token总 token单位时间得分每成功任务成本
HLE (w/o tools)*1.02−2%−10%+5%0.98×+5%
HLE (w/ tools)0.88+14%−22%+3%1.15×+2%
Terminal Bench 2.10.79+27%−28%−3%1.29×−5%
DeepSWE0.78+28%−28%−3%1.34×−7%
Toolathlon-Verified0.86+16%−25%+2%1.19×+0%
AutomationBench (Public)0.76+32%−31%−5%1.41×−12%

* HLE (w/o tools) 的 τ=1.02 是有意为正(即变慢):单轮任务上技能条目是净开销。

Contributors

Tiger3807861189

21 commits