naïvelabs
[Research to Advance Agent Infrastructure]
Naïve is a research lab with one objective function: multiplying intelligence per token.
[Benchmarks]
Sandbox
View moreYou pay for an isolate, not a machine.
| System | Cold start | RAM / agent | Idle agent |
|---|---|---|---|
| Naïve | 2.79msisolated-vm | 1.2MBisolated-vm | KBs of statestorage only |
| E2B | <200ms | 512MBminimum | Paused VMper GiB |
| Modal | ~1s | 128MBminimum | Reserved |
| Morph | — | — | Reserved VM |
| Cloudflare | 1–3s | 256MBminimum | Memory reserved |
Agent Brain
View moreRecall no other memory system reaches.
| System | Single-hop | Multi-hop |
|---|---|---|
| Naïve Brain | 91% | 57% |
| CAR (prior SOTA) | 78% | 30.2% |
| HippoRAG-v2 | 54% | ≤7% |
| MemGPT / Cognee | 28% | ≤7% |
| Mem0 / Contriever | 18% | ≤7% |
| Zep / Graphiti | 7% | ≤7% |
Inference
View moreFrontier models waste money searching.
| System | With Scout | $/solved | Δ cost |
|---|---|---|---|
| Claude Opus 4.8 | 36/5072% | $1.23was $1.75 | −29.8% |
| Claude Opus 5 | 37/5074% | $1.27was $1.61 | −21.0% |
| GPT-5.6 Sol | 34/5068% | $0.82was $1.05 | −21.9% |
| GLM-5.2 | 27/5054% | $0.60was $1.25 | −52.0% |
Agent Workforce
View moreCost of hosting 1 million agents per month
| System | 1M agents / mo | Idle |
|---|---|---|
| Naïve Vetta · serverless | ~$44–60kmodelled | Storage onlyno compute |
| Hosting Hermes / Eve / OpenClaw / Pi | ~$740k–1.5Mmodelled | Billed 24/7 |
| VM per tenant · AWS | ~$14Malways-on | Billed 24/7 |
Feature coverage · 24 capabilities
Naïve14/24
Hermes12/24
OpenClaw14/24
Paperclip8/24
Eve9/24
Pi3/24
Agent Loop Efficiency
Same intelligence, 30% more efficient.
| System | SWE-bench solve | $ / solved | Δ cost |
|---|---|---|---|
| Naïve Vetta | 69.83% | $1.19 | −30.41% |
| Claude Code | 70.50% | $1.71 | baseline |
| Claude Managed Agents | 68.24% | $1.62 | −5.26% |
| Hosted Hermes | 68.97% | $1.89 | +10.53% |
[Papers]
| Date | Title | Category | Authors |
|---|---|---|---|
| 8/4/2026 | Building Towards the Most Efficient Agent Loop Infrastructure | Infrastructure, Benchmarks | Naïve Research |
| 7/26/2026 | Sandboxes Without Machines: A V8-Isolate Runtime for Massive Fleets of Resident AgentsComing soon | Sandbox, Runtime | Naïve Research |
| 7/26/2026 | Serverless Agents Without Machines: Scale-to-Zero for Fleets of Mostly-Idle AgentsComing soon | Serverless, Orchestration | Naïve Research |
| 7/27/2026 | Both Engines, One Harness: A Self-Run Head-to-Head Against mem0's Open-Source ServerComing soon | Memory, Benchmarks | Naïve Research |
[Careers]
| Position | Location | |
|---|---|---|
| Member of Technical Staff, Agent Memory | San Francisco · On-site | Apply on YC |
| Member of Technical Staff, Inference | San Francisco · On-site | Apply on YC |
| Member of Technical Staff, Agent Orchestration | San Francisco · On-site | Apply on YC |