Eval and Treat CPT Code Referral

DeepCode: Open Agentic Coding

We evaluate DeepCode on the PaperBench benchmark (released by OpenAI), a rigorous testbed requiring AI agents to independently reproduce 20 ICML 2024 papers from scratch. The benchmark comprises 8,316 ...

GitHub

CATArena: Engineering-Level Tournament Evaluation Platform for LLM-Driven Code Agents

CATArena (Code Agent Tournament Arena) is an open-ended environment where LLMs write executable code agents to battle each other and then learn from each other. CATArena is an engineering-level ...

PC World

I started ‘vibe coding’ my own apps with AI and I’m utterly loving it

On February 2nd, 2025, computer scientist and OpenAI co-founder Andrej Karpathy made a flippant tweet that launched a new phrase into the internet’s collective consciousness. He posted that he’d ...

Some results have been hidden because they may be inaccessible to you

Show inaccessible results

DeepCode: Open Agentic Coding

CATArena: Engineering-Level Tournament Evaluation Platform for LLM-Driven Code Agents

I started ‘vibe coding’ my own apps with AI and I’m utterly loving it

Trending now