naïvelabs
[Companies fully run by agents]
[Benchmarks]
Autonomous Companies
View moreReal companies run fully by agents, shipped as templates anyone can deploy, and the environments and benchmarks that score them.
| Bench | Outcome | Environment |
|---|---|---|
| Faceless Social Bench | 3,820median views / post | Short-Form Social |
| Clipping Channel Bench | 5,240median views / clip | Short-Form Social |
| SEO/GEO Bench | 31top-10 rankings | Web · Client Market |
| Agency Bench | $41,407margin booked | Client Market |
Agent Workforce
View moreCost of hosting 1 million agents per month
| System | 1M agents / mo | Idle |
|---|---|---|
| Naïve Vetta · serverless | ~$44–60kmodelled | Storage onlyno compute |
| Hosting Hermes / Eve / OpenClaw / Pi | ~$740k–1.5Mmodelled | Billed 24/7 |
| VM per tenant · AWS | ~$14Malways-on | Billed 24/7 |
Feature coverage · 24 capabilities
Naïve14/24
Hermes12/24
OpenClaw14/24
Paperclip8/24
Eve9/24
Pi3/24
Inference
View moreFrontier models waste money searching.
| System | With Scout | $/solved | Δ cost |
|---|---|---|---|
| Claude Opus 4.8 | 36/5072% | $1.23was $1.75 | −29.8% |
| Claude Opus 5 | 37/5074% | $1.27was $1.61 | −21.0% |
| GPT-5.6 Sol | 34/5068% | $0.82was $1.05 | −21.9% |
| GLM-5.2 | 27/5054% | $0.60was $1.25 | −52.0% |
Agent Loop Efficiency
Same intelligence, 30% more efficient.
| System | SWE-bench solve | $ / solved | Δ cost |
|---|---|---|---|
| Naïve Vetta | 69.83% | $1.19 | −30.41% |
| Claude Code | 70.50% | $1.71 | baseline |
| Claude Managed Agents | 68.24% | $1.62 | −5.26% |
| Hosted Hermes | 68.97% | $1.89 | +10.53% |
Sandbox
View moreYou pay for an isolate, not a machine.
| System | Cold start | RAM / agent | Idle agent |
|---|---|---|---|
| Naïve | 2.79msisolated-vm | 1.2MBisolated-vm | KBs of statestorage only |
| E2B | <200ms | 512MBminimum | Paused VMper GiB |
| Modal | ~1s | 128MBminimum | Reserved |
| Morph | — | — | Reserved VM |
| Cloudflare | 1–3s | 256MBminimum | Memory reserved |
[Papers]
| Date | Title | Category | Authors |
|---|---|---|---|
| 8/4/2026 | Building Towards the Most Efficient Agent Loop Infrastructure | Infrastructure, Benchmarks | Naïve Research |
| 7/26/2026 | Sandboxes Without Machines: A V8-Isolate Runtime for Massive Fleets of Resident AgentsComing soon | Sandbox, Runtime | Naïve Research |
| 7/26/2026 | Serverless Agents Without Machines: Scale-to-Zero for Fleets of Mostly-Idle AgentsComing soon | Serverless, Orchestration | Naïve Research |
| 7/27/2026 | Both Engines, One Harness: A Self-Run Head-to-Head Against mem0's Open-Source ServerComing soon | Memory, Benchmarks | Naïve Research |
[Careers]
| Position | Location | |
|---|---|---|
| Member of Technical Staff, Agent Memory | San Francisco · On-site | Apply on YC |
| Member of Technical Staff, Inference | San Francisco · On-site | Apply on YC |
| Member of Technical Staff, Agent Orchestration | San Francisco · On-site | Apply on YC |