Skip to content
#

independent-benchmark

Here are 5 public repositories matching this topic...

Open, independent benchmark on web search APIs for AI agents - factual lookup on 300 company news questions, ranked by cost per 1,000 correct answers. TinyFish, Parallel, Perplexity, Linkup, Firecrawl, Brave, You, Exa, Tavily, Google SERP.

  • Updated Sep 4, 2026
  • Python

Open, independent benchmark on LLM inference providers: same GLM 5.3 Flash, 600 paired requests each. Latency, tokens per second and task success. Baseten, DeepInfra, Fireworks AI, Modal, Nebius, Novita AI, Parasail, Telnyx, Together AI, Z.AI. Python runner and public data.

  • Updated Sep 5, 2026
  • Python

Open, independent benchmark on company news APIs for AI agents - web search APIs vs dedicated news indexes on 300 company news questions, ranked by cost per 1,000 correct answers. Exa, Parallel, Perplexity, Brave, Firecrawl, Linkup, You, TinyFish, Tavily, Google SERP, Seltz, PredictLeads, Autobound, Datahyena.

  • Updated Sep 5, 2026
  • Python

Open, independent benchmark of web search APIs for coding agents: Exa, Parallel, Perplexity, Firecrawl, Tavily, Linkup, Brave, You, TinyFish. Scored on grounded task completion against held-out enterprise docs tickets. Search-only and search+fetch boards, model held constant.

  • Updated Sep 9, 2026
  • Python

Open, independent benchmark on web search APIs for deep research agents - unlike BrowseComp, this is a benchmark on real user workflows that are not memorized by models. Exa, Parallel, Perplexity, Linkup, Tavily, Firecrawl, Brave, You, TinyFish, Seltz, Google SERP.

  • Updated Sep 4, 2026
  • Python

Add this topic to your repo

To associate your repository with the independent-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more