← 返回名录条目信息来自其公开主页与公开发布内容。按本站规范,页面不展示任何联系方式。 需要更正或删除?通过收录与更正通道提交,24 小时内处理。
代表作品 / 项目
- opencompass⭐ 7492OpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100+ datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.
- MathBench⭐ 118[ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset
- CompassVerifier⭐ 71[EMNLP 2025] CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
- GPassK⭐ 33[ACL 2025] Are Your LLMs Capable of Stable Reasoning?
- RePro⭐ 14[ICLR 2026] Rectifying LLM Thought From Lens of Optimization