Preprint / Version 1

AI Agents Need Operating Systems, Not Just Better Models

Why the Next Frontier of AI Reliability Is a Software Engineering Problem, Not a Modeling One

##article.authors##

DOI:

https://doi.org/10.31224/7970

Keywords:

AI agents, agent runtime, software reliability, runtime verification, provenance, saga pattern, multi-agent systems, distributed systems

Abstract

Production failures in AI agent deployments are, in most cases, software engineering failures caused by the absence of principled runtime infrastructure — not model capability failures. This article argues that today's agent frameworks solve composition problems but leave the runtime layer untouched: process isolation, resource governance, structured error handling, durable state, and provenance capture. Drawing on operating systems design, the saga pattern, runtime verification, and transaction theory, it lays out six reliability properties a mature agent execution environment needs, walks through where frameworks such as LangChain and orchestration engines such as Temporal and Airflow fall short, and closes with concrete engineering steps teams can take today without waiting for a dedicated agent runtime to exist.

Downloads

Download data is not yet available.

Downloads

Posted

2026-08-17