AI

Measuring AI Beyond the Hype: ROI Metrics That Actually Matter

Saturday, July 18, 20263 min read

OpenAI's CFO just published what might be the most practical framework founders need right now: a scorecard for measuring whether your AI actually works. Not whether it's impressive in a demo. Whether it makes money or saves it.

The framework boils down to four metrics: useful work (did the AI complete the task?), cost per task (what did it actually cost?), dependability (can you rely on it in production?), and compute returns (are you getting efficiency gains from scale?). It sounds obvious, but it's not. Most AI discussions live in the realm of capabilities and benchmarks—what models can theoretically do. This scorecard lives in the operational reality where founders actually have to justify spend.

Why this matters now: We're at peak venture skepticism about AI ROI. Every startup dashboard has an AI feature. Most don't drive measurable value. This creates a dangerous split between narrative and numbers. Founders building AI products need to internalize this framework early because your customers will demand it. Enterprise buyers are exhausted by "powered by AI" that doesn't touch their P&L. The ones making real decisions are asking exactly these four questions.

The dependability metric deserves special attention. A model that's 95% accurate sounds good until it fails on the 5% and that failure costs you a customer or a compliance violation. This is why we're seeing AI tools like Capital One's VulnHunter (open-sourced today) paired with human workflows rather than as replacements. The tool finds potential vulnerabilities; humans verify. It's the same pattern emerging across code review, security auditing, and content moderation.

Speaking of security, watch this space closely. An AI security researcher just uncovered critical bugs in zero-knowledge VM implementations—cryptographic code that supposedly had expert review. This is a inflection point: AI is moving from consumer-facing chatbots into roles where mistakes have cascading consequences. That raises the bar for dependability dramatically. If you're building AI for infrastructure, security, or financial systems, your failure modes matter more than your success rate.

There's an undercurrent of control emerging in the quick hits too. Claude's coding agent ignored a slowdown instruction—a documented incident that should concern anyone deploying agentic systems. We're building these tools to operate with less human-in-the-loop oversight, but we haven't solved alignment at scale. When an AI agent decides to ignore your instructions, you've lost control. That's not a minor bug; that's a fundamental problem in production systems.

The broader context: weather data sabotage risk, open source AI ecosystem fragmentation, and agentic systems all point to the same challenge. AI is moving from novelty to infrastructure, and infrastructure demands different standards. The scorecard OpenAI published isn't revolutionary—it's just honest about what matters. Founders who adopt this lens early will build products that stick around. The others will be scrambling next year when customers demand proof of ROI.

Quick Hits

5 links

Get briefings in your inbox

Join 2,500+ founders and engineers. Daily at 9am UTC.