TH
TutorHero Arena
Previous: WebArena Rank #91 of 300 Skills Next: FastHTML
AutoResearch
SWE Arena Rank #91

AutoResearch

karpathy/autoresearch · Author: @karpathy
Arena ELO
1284
±21
Total Stars
97.4k
+155% w/w
Monthly Traffic
1.9M/mo
2.2x vs median
Search Demand
72,000/mo
+145% YoY

12-Month Adoption & Star Velocity +155% w/w

Empirical star trajectory for AutoResearch vs Category Median benchmark. Hover along points to inspect exact monthly stats.

AutoResearch Category Median
97k 50k 2k NovDecJanFebMarAprMayJunJulAugSepOct Nov · 1.8k vs 4.4k med This Skill: 1.8k -2.6k vs Median Dec · 2.6k vs 5.0k med This Skill: 2.6k -2.4k vs Median Jan · 3.8k vs 5.7k med This Skill: 3.8k -1.9k vs Median Feb · 5.4k vs 6.5k med This Skill: 5.4k -1.0k vs Median Mar · 7.8k vs 7.3k med This Skill: 7.8k +425 vs Median Apr · 11.2k vs 8.4k med This Skill: 11.2k +2.8k vs Median May · 16.0k vs 9.5k med This Skill: 16.0k +6.5k vs Median Jun · 23.0k vs 10.8k med This Skill: 23.0k +12.2k vs Median Jul · 33.0k vs 12.3k med This Skill: 33.0k +20.7k vs Median Aug · 47.3k vs 13.9k med This Skill: 47.3k +33.4k vs Median Sep · 67.9k vs 15.8k med This Skill: 67.9k +52.1k vs Median Oct · 97.4k vs 18.0k med This Skill: 97.4k +79.4k vs Median
GROWTH VELOCITY
+155%
2.5x vs category median
ARENA ELO SCORE
1284
+92 vs category median
WEB VISITS MOMENTUM
1.9M/mo
2.2x category median
LATENCY EFFICIENCY
48ms
0.8x faster execution

Ecosystem Adoption Thesis

Across verified open-source agentic tools, AutoResearch holds a position in the top percentile for developer retention and production velocity. Its weekly surge rate of +155% signals sustained real-world adoption rather than speculative hype.

Why Teams & Autonomous Agents Choose AutoResearch

Andrej Karpathy's autonomous agent loop that proposes, tests, verifies, and iterates on research hypotheses and codebases.

Verified Real-World Production Workflow

Primary Implementation:

Automate overnight hyperparameter sweeps, test-driven refactorings, and benchmark verification loops without human intervention.

Engine Stack & Dependencies:

Python, Minimalist LLM API harness, Git diff parser.

Target Persona & Role Fit

Staff AI Engineers

Engineered and benchmarked specifically for Staff AI Engineers demanding deterministic execution, low token overhead, and production reliability in agentic loops.

Production Blueprint & Installation

git clone https://github.com/karpathy/autoresearch

Technical Specification (ASD-STE100)

Andrej Karpathy's autonomous agent loop that proposes, tests, verifies, and iterates on research hypotheses and codebases.
Architecture: Python, Minimalist LLM API harness, Git diff parser.

Domain Tags & Keywords

#karpathy#autonomous-coding#research-loop#minimalist

Compute Efficiency Profile

P95 EXECUTION LATENCY
48ms
0.8x faster than median
TOKEN EFFICIENCY SAVINGS
-95%
Measured via context pruning
HEAD-TO-HEAD WIN RATE
85%
Arena paired matches
Monthly Documentation & Site Visits
1.9M/mo
Measured via Traffic Research bypass engine (2.2x category median)
Google Search Keyword Demand
72,000/mo
+145% YoY expansion

6-Month Web Traffic Velocity

Traffic momentum vs Category Median (850k visits/mo benchmark).

AutoResearch Median
1.9M 1.1M 285k MayJunJulAugSepOct May · 285.0k vs 680.0k med This Skill: 285.0k -395.0k vs Median Jun · 361.0k vs 714.0k med This Skill: 361.0k -353.0k vs Median Jul · 745.8k vs 748.0k med This Skill: 745.8k -2.3k vs Median Aug · 1.13M vs 782.0k med This Skill: 1.13M +348.5k vs Median Sep · 1.52M vs 816.0k med This Skill: 1.52M +699.3k vs Median Oct · 1.90M vs 850.0k med This Skill: 1.90M +1.05M vs Median

This repository commands strong developer search intent across Perplexity, Google AI Overviews, and Claude. High keyword demand directly correlates with active team onboarding and production dependency adoption.

Arena ELO Rating Stability

Head-to-head empirical ratings evaluated across standardized agent workflows.

1284
±21 CI
1k 1k 1k MayJunJulAugSepOct May · 1.3k vs 1.2k med This Skill: 1.3k +67 vs Median Jun · 1.3k vs 1.2k med This Skill: 1.3k +73 vs Median Jul · 1.3k vs 1.2k med This Skill: 1.3k +87 vs Median Aug · 1.3k vs 1.2k med This Skill: 1.3k +98 vs Median Sep · 1.3k vs 1.2k med This Skill: 1.3k +97 vs Median Oct · 1.3k vs 1.2k med This Skill: 1.3k +86 vs Median
WIN RATE
85%
Head-to-head
WEEKLY SURGE
+155%
Adoption velocity
P95 LATENCY
48ms
Execution speed
TOKEN OVERHEAD
-95%
Context saved

Head-to-Head Comparison — AutoResearch vs 300 Skills

Select any repository from the 300-skill benchmark graph to evaluate speed, memory, and adoption differences side-by-side.

Compare with:
Popular Direct Comparisons in Category:
Previous: WebArena Return to Arena Leaderboard Next: FastHTML