Skill Claude CodeCodex
A question-creation tool for the CAC evaluation system, which tests whether software performs expected tasks.
20 1mo ago A 81 tokens
AGPL-3.0
Benchmark for Evaluating the Performance of Natural Language Models and Agents / 自然语言模型 & Agent性能评测基准
Skill Claude CodeCodex
A question-creation tool for the CAC evaluation system, which tests whether software performs expected tasks.