A comprehensive evaluation framework for Large Language Models (LLMs), providing extensive assessments across three key dimensions: general capabilities, safety, and robustness. The framework includes diverse benchmarks and supports both API-based and local models with distributed evaluation capabilities.
-
Updated
Jun 27, 2025 - Python