Framework for evaluating and improving agents
-
Updated
Aug 17, 2026 - Python
Framework for evaluating and improving agents
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
A Universal Platform for Training and Evaluation of Mobile Interaction
A graphical interface for reinforcement learning and gym-based environments.
Interoperating between (Deep) Reiforcement Learning libraries
Gymnasium-style API standard for RL environment creation in JAX
Create new gridworld gym environments easily
Workspace manager for coding agents. Interactively solve and develop Harbor tasks.
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
Comprehensive AI agent evaluation platform — searchable benchmark catalog, comparison matrices, automated scanner, interactive dashboards, and community-curated best practices for LLM evaluation.
Outcome-verified agent trajectories, benchmarks, and RL environments — with a live leaderboard and a CI gate for your agents. Offline-first, MIT.
Agent-evaluation environments: planted-truth worlds, ungameable graders, calibrated difficulty
Sound error bounds, symbolic GPU safety checks and targeted falsification for Triton kernels. 88% of planted bugs pass the standard fixed-shape allclose test; Litmus catches 100% with 0% false positives, on CPU.
Foundry Lite: a public runnable sample of Veyl’s local environment harness for software-engineering agent evals.
A lightweight, open-source framework that turns historical GitHub pull requests into reproducible, verifiable software-engineering tasks for training and evaluating coding agents.
Open-source RL environments and evals for the capabilities we want AI to have. Judge-free scoring, mandatory baselines, and an enforced defensive-asymmetry gate.
Claude Code Agent Skills for building, red-teaming and tuning agentic RL evaluation environments — a four-skill pattern (guardian, validation-debugger, score-tuner, iteration-loop) plus a 24-point adversarial reviewer.
Open-source SDK (Apache-2.0): RL environments, conformal calibration, a TRL-compatible reward function, the Lean 4 formal track, and the verifiable/vlabs CLI.
Add a description, image, and links to the rl-environments topic page so that developers can more easily learn about it.
To associate your repository with the rl-environments topic, visit your repo's landing page and select "manage topics."