Quesma
Making AI agents production-ready through independent evaluation and training.
Pinned Loading
Repositories
Showing 10 of 23 repositories
- awesome-ai-tokenomics Public
A curated list of tools, benchmarks, papers, and copy-paste configs for AI token costs: what tokens cost, where they get wasted, and how to cut the bill.
- terminal-bench-science Public Forked from harbor-framework/terminal-bench-science
Terminal Bench for Science
- BinaryAudit Public
An open-source benchmark for evaluating AI agents' ability to find backdoors hidden in compiled binaries.
- trival-prompt-bench Public
- CompileBench Public
Benchmark of LLMs on real open-source projects against dependency hell, legacy toolchains, and complex build systems.
Top languages
Loading…
Most used topics
Loading…