Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 39 public sources, with linked evidence and daily updates.
-
Updated
Oct 7, 2026 - Python
Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 39 public sources, with linked evidence and daily updates.
AI4AI Survey: can AI reliably improve AI? 223 papers on long-horizon agents, benchmarks, harness design, and recursive self-improvement · updated weekly
Official companion repository for our survey "A Survey of the OpenClaw Ecosystem: From Platform Extensibility to Constraint Design" — a curated collection of papers, benchmarks, security reports, datasets, and tools for the OpenClaw AI agent ecosystem.
Strict type, range, and semantic key validator for agent tool call invocations
Strict type, range, and semantic key validator for agent tool call invocations
Step-by-step agent trajectory evaluator comparing execution traces against golden tool call sequences
Step-by-step agent trajectory evaluator comparing execution traces against golden tool call sequences
[EMNLP 2026] Messier is a cross-benchmark corpus of AI agent evaluations, linking tasks and results with models, scaffolds, verifiers, scoring rules, and execution trajectories.
An evidence-hostile, container-isolated benchmark and behavior analysis platform for long-horizon AI software agents.
Anti-cheating audit and isolated evaluation sandbox for AI agent benchmarks: detect the 7 vulnerability patterns, grade on the Agent-Eval Checklist, red-team with zero-capability agents. Zero dependencies.
A curated list of frontier AI evaluation benchmarks, long-horizon agent test suites, terminal sandboxes, RLVR math verifiers, and red-teaming datasets.
Agent evaluation, verifier engineering, and cloud-native systems
A free Claude Code skill to test coding agents on the same tasks; compare pass rate, cost, and time—built on affaan-m/ECC.
Generated by Codex. Read with caution.
Deterministic release decisions from agent benchmark evidence.
用同一套提示词评测多模型 Agent 从零实现完整可玩游戏的能力:固定技术栈、固定评分标准、按模型归档对比。| A standardized prompt for benchmarking LLM agents on building a complete game from scratch — fixed tech stack, fixed rubric, per-model results.
Inventory expiry regression tests for AI agent business simulators. Detect stock-age resets and expired sales in pinned Vending Bench and RetailBench implementations; no LLM calls.
To associate your repository with the agent-benchmarks topic, visit your repo's landing page and select "manage topics."