THE AI STARTUP REPORT

Field guide

Enterprise agents have entered their software era

The hard problem is no longer making an agent look capable. It is defining, testing, governing, and improving one after it reaches production.

01

A demo is an argument, not a deployment

The first generation of agent demonstrations established that models could plan, call tools, and complete sequences of work. That was an important proof, but it created a misleading picture of the product challenge. A successful demo begins with a selected task, a prepared environment, and an observer who knows what should happen. Production begins when the task is ambiguous, the data is incomplete, the policy changed yesterday, and the user assumes the software will behave correctly without supervision.

Enterprise agents are therefore becoming less like magical coworkers and more like software systems with a new interface. They need a definition, an owner, access controls, test cases, release management, monitoring, and a process for handling failure. The companies building those surrounding capabilities may hold a more durable position than products that compete only on the apparent intelligence of a single interaction.

02

The lifecycle is becoming visible

Glean recently described an Agent Development Lifecycle built around defining, building, launching, governing, and improving agents. The framework is revealing because it imports the discipline of software development into a category that has often been presented as no-code spontaneity. Natural language can make an agent easier to create, but it does not eliminate the need to know what the system is supposed to do or how its behavior will be evaluated.

A mature agent platform needs a sandbox, traces, version history, policies, and measurements connected to a business outcome. It also needs a way to discover redundant agents before every team builds its own slightly different version. The risk of agent sprawl is not only wasted spending. It is inconsistent behavior, duplicated access to sensitive systems, and a growing set of automations no one is confident enough to change.

03

Business users and technical teams need the same surface

Decagon's product is built around a productive tension. Customer-operations teams need direct control over procedures and improvements, while technical teams need visibility into integrations, guardrails, and versions. If every policy adjustment requires a vendor ticket, the agent cannot keep pace with the business. If every team can change behavior without controls, the organization inherits a different problem.

Natural-language operating procedures may become a bridge between those groups. They make logic legible to the people closest to the customer while creating an artifact that can be tested and reviewed. The important product work lies beneath the friendly interface: resolving ambiguity, showing the effect of a change, and preventing an instruction that sounds reasonable from producing unsafe behavior in an unexpected situation.

04

Vertical agents carry a higher burden

Harvey demonstrates why domain-specific agents require more than general tool use. Legal and professional work depends on sources, matter context, confidentiality, and an accountable reviewer. An agent can accelerate research or document analysis, but the professional must remain able to understand where the output came from and what assumptions shaped it. Fluency is not a substitute for authority.

The same principle applies in healthcare, finance, and other consequential fields. The agent's product surface must include provenance and escalation, not merely an answer. Vertical companies can turn these constraints into an advantage because they can design evaluations, workflows, and controls around a narrower domain. General platforms will remain powerful, but specialized systems can earn trust by understanding what failure means in a particular profession.

05

Observability becomes a product feature

Traditional software is usually deterministic enough that teams monitor availability, latency, and errors. Agents introduce a softer class of failure: the system completes the task, but does so poorly, uses the wrong evidence, takes an unnecessary path, or produces an outcome that conflicts with the user's intent. Monitoring must therefore include conversation quality, task success, policy adherence, tool selection, and the reasons people override the result.

This is why simulation and review are moving toward the center of agent platforms. Teams need to replay difficult cases, compare versions, and learn from the conversations where a person had to intervene. Over time, the evaluation set becomes a form of institutional knowledge—a record of what the organization considers correct, acceptable, and risky.

06

The winners will make restraint visible

The public imagination focuses on what an agent can do. Enterprise adoption will often depend on what it refuses to do. A dependable agent recognizes missing information, preserves permissions, respects an approval boundary, and transfers a case with enough context for a person to continue. Restraint is difficult to market in a short demonstration, but it is one of the strongest signs that a system is ready for consequential work.

The next phase of the agent market will be measured by governed outcomes rather than autonomous spectacle. Products will still compete on model quality and user experience, but the durable layer will include the unglamorous machinery of software: ownership, testing, access, observability, and change management. Agents are not escaping the requirements of enterprise software. They are expanding them.

Primary sources

Glean Agent Development LifecycleDecagon product overviewSierra company overviewHarvey company overview