Lenny's Podcast · Published 2026-08-30

AI’s third era: the rise of persistent AI coworkers | Tara Seshan (OpenAI’s product lead)

Open on YouTube ↗

Summary

Overview

  • Speaker: Tara Seshan
  • Channel: Lenny's Podcast
  • Main topic: The evolution of AI products from chat interfaces to persistent AI coworkers in knowledge work.
  • Purpose: To provide deep product management insights on building AI tools, navigating organizational culture at OpenAI, and understanding the future of knowledge work in the age of AI agents. Tara Seshan, Product Lead for Codex and ChatGPT Work at OpenAI, discusses the three eras of AI products: chat, agents, and persistent coworkers. She shares product management frameworks, the shift from theoretical to empirical building, working at OpenAI versus traditional companies, and how knowledge workers must adapt to a world where AI handles execution and humans focus on strategy and direction.

Topic Map

The Three Eras of AI Products

  • Explanation: AI products are evolving through three distinct phases: chat interfaces (era 1), task-oriented agents (era 2), and persistent AI coworkers (era 3).
  • Key claims:
    • Era 1 is chat-based interfaces.
    • Era 2 is working with autonomous agents.
    • Era 3 is persistent AI coworkers who get things done alongside you.
  • Examples:
    • Basic chatbots
    • Autonomous agentic workflows
    • Persistent AI co-workers embedded in daily work
  • Terminology:
    • chat era
    • agent era
    • coworker era
  • Why it matters: Helps product builders understand where AI product utility is heading and avoid building for outdated paradigms.

Building for the Next 2 to 3 Months

  • Explanation: In an exponentially fast-moving AI landscape, building for today or projecting one year out leads to failure; teams must build for the immediate 2-3 month capability horizon.
  • Key claims:
    • Building for current models fails.
    • Building for where models will be in a year fails.
    • The only viable window for product planning is 2 to 3 months out.
  • Examples:
    • Iterating rapidly on Codex and ChatGPT Work features based on current frontier capabilities.
  • Terminology:
    • model capability horizon
    • empirical building
  • Why it matters: Prevents wasted engineering effort on products designed for obsolete or overly futuristic model capabilities.

Prolific and Empirical vs. Academic and Theoretical

  • Explanation: Being prolific and empirical—getting things into users' hands as fast as possible—is far more valuable in AI product management than writing long theoretical reasoning docs.
  • Key claims:
    • Empirical testing beats academic planning.
    • Rapid prototyping with users accelerates product-market fit.
  • Examples:
    • Releasing features quickly and iterating based on real user telemetry rather than internal debate.
  • Terminology:
    • empirical product development
    • prolific iteration
  • Why it matters: Changes how product teams operate in fast-moving tech environments.

Founders-Led vs. Founders Led Culture at OpenAI

  • Explanation: While most companies are founder-led, OpenAI is 'founders-led,' where top-down direction is minimal and individuals in every area act as founders of their product domains.
  • Key claims:
    • OpenAI is founders-led, giving product managers extreme autonomy.
    • Distance between the product creator and the market is extremely thin.
  • Examples:
    • Product managers operating with founder-like velocity and ownership.
  • Terminology:
    • founders-led
    • product autonomy
  • Why it matters: Explains the high velocity and output of teams operating at the frontier of AI research.

The Shift from Rowing to Steering in Knowledge Work

  • Explanation: As AI agents take over execution (rowing), the human role in knowledge work shifts to direction and orchestration (steering).
  • Key claims:
    • Knowledge work is shifting from execution to steering.
    • AI handles tactical execution while humans define core hypotheses and goals.
  • Examples:
    • Using AI agents to generate financial models or write code while humans set direction and review outcomes.
  • Terminology:
    • steering vs rowing
    • agentic delegation
  • Why it matters: Redefines professional roles and productivity expectations for knowledge workers.

Key Points

Empirical testing over theoretical planning

  • Explanation: In AI, you cannot theorize a long-term product roadmap because capabilities shift too fast; you must build, test with users, and iterate rapidly.
  • Evidence: Tara's experience transitioning from academic-heavy environments to OpenAI's rapid empirical approach.
  • Practical implication: Ship prototypes to users immediately rather than writing extensive product specification documents.

The PM's core job is hypothesis definition and rapid testing

  • Explanation: With AI tools accelerating execution, the product manager's primary role is sharpening the core hypothesis and testing it against reality as fast as possible.
  • Evidence: Discussion on how the distance between PM and market is compressed at OpenAI.
  • Practical implication: Focus less on operational overhead and more on defining the exact question to test.

Mainlining the product

  • Explanation: Teams building AI products must use their own products constantly and make them central to their daily operating cadence.
  • Evidence: OpenAI internal culture of using ChatGPT Work and Codex continuously.
  • Practical implication: Dogfooding internal AI tools aggressively accelerates product refinement.

Frameworks, Models & Processes

The Three-Month Horizon Framework

  • How it works: Plan and execute product features based strictly on model capabilities expected 2 to 3 months out, ignoring both current limitations and distant 1-year projections.
  • Components:
    • Current model baseline assessment
    • Short-term research roadmap integration
    • Rapid user feedback loop
  • When to use: When building applications on top of rapidly evolving foundational AI models.

Writing as Thinking vs. Writing as Reporting

  • How it works: Distinguish between writing to formulate your own thoughts (which cannot be automated) and writing to report status or summarize (which should be automated).
  • Components:
    • Personal drafting and outlining for thinking
    • Delegating summaries and status reports to AI
  • When to use: When managing knowledge work tasks and deciding what to delegate to AI agents.

Examples & Case Studies

Transitioning from Stripe to OpenAI.

  • Illustrates: Moving from an established, scaled hyper-growth company to an extreme frontier research lab.
  • Lesson: Founders-led cultures require extreme personal ownership and rapid adaptation to unstructured environments.

Building and using Sites within Codex/ChatGPT workflows.

  • Illustrates: How AI can instantly generate functional digital artifacts from simple prompts.
  • Lesson: Products that eliminate the gap between thought and artifact creation radically change knowledge work.

Actionable Takeaways

  • Immediate:
    • Adopt a 2 to 3-month product planning horizon.
    • Embrace empirical prototyping over lengthy specification documents.
    • Use AI tools to handle execution while retaining ownership of core hypotheses.
  • Strategic:
    • Prepare for knowledge work to transition entirely from rowing (execution) to steering (direction).
    • Foster a founders-led culture where individuals have high autonomy and direct connection to users.
    • Mainline your own AI products internally to ensure deep product intuition.
  • Questions to investigate:
    • How will enterprise software architectures adapt to persistent AI agents with broad data access?
    • What new evaluation metrics replace traditional product-market fit when AI capabilities compound weekly?

Claims Worth Verifying

  • OpenAI operates with a founders-led culture where top-down direction is extremely limited. (organizational structure)
  • ChatGPT Work is the fastest-growing AI product for knowledge workers. (market statistic)

Notable Quotes

"you fail if you build for where the models are now, you fail if you build for where you think the models will be in a year, both outcomes are equally wrong, the only way to build is two to three months." "prolific and empirical is way more important than being academic or theoretical." "it's not like real estate, you don't like put money in and get value out, it is a little bit more like filmmaking."

Compressed Summary

  • AI products are evolving from chat to agents to persistent coworkers.
  • Product teams must build within a strict 2-3 month capability horizon.
  • Empirical testing and rapid prototyping beat theoretical planning.
  • Knowledge work is shifting from execution (rowing) to direction (steering).
  • Keywords: artificial intelligence, product management, openai, knowledge work, agents
  • Core insight: In the era of AI coworkers, knowledge work shifts from execution to steering, requiring teams to build empirically on a compressed 2-3 month horizon.

Core insights

6
Architecturehigh noveltymoderate evidence

The paradigm shift is not just from chat to agents, but from episodic agents to persistent AI coworkers: systems that are embedded in daily work and act alongside humans over time rather than being invoked for a task and disappearing.

Why it matters

This changes where product and system boundaries belong. An agentic architecture designed for persistent coworkers needs long-lived state, shared workspace context, interruption/background work handling, and recovery semantics beyond a single request/response loop.

Generalization

When designing an AI product, decide whether it is a tool that executes on command or a coworker with ongoing presence, context, and ownership over work over time.

AI products are evolving through three distinct phases: chat interfaces (era 1), task-oriented agents (era 2), and persistent AI coworkers (era 3).
Open source video
Era 3 is persistent AI coworkers who get things done alongside you.
Open source video
Practicemedium noveltymoderate evidence

Products built on rapidly evolving frontier models should not be designed for today's model or for a speculative one-year future; the effective planning window is about 2 to 3 months out.

Why it matters

Engineering roadmaps, feature selection, and architecture decisions depend heavily on which model capabilities are assumed. Too short a horizon produces work that is obsolete; too long a horizon produces work optimized for capabilities that may never materialize as imagined.

Generalization

For AI application teams, time-bound capability assumptions to a short concrete horizon and isolate them from more stable parts of the system.

Building for current models fails.
Open source video
Building for where models will be in a year fails.
Open source video
The only viable window for product planning is 2 to 3 months out.
Open source video
Mental Modelmedium noveltymoderate evidence

Persistent coworker delegation requires separating rowing from steering: AI performs tactical execution while humans own hypotheses, goals, and outcome review.

Why it matters

This is a concrete control-plane design principle. Agent orchestration should expose the agent's plan and assumptions so a human can steer it, rather than treating the agent as a black-box doer.

Generalization

Any delegation system needs a clearly defined owner of the objective and a review loop over the agent's execution artifacts.

Knowledge work is shifting from execution to steering.
Open source video
AI handles tactical execution while humans define core hypotheses and goals.
Open source video
Mechanismhigh noveltymoderate evidence

A useful criterion for what to delegate to an AI is whether the writing is thinking or reporting: reporting and summarization should be automated, while drafting and outlining that form one's own thoughts should not.

Why it matters

Naive delegation of all writing to AI can automate away the cognitive work that makes knowledge work valuable. This criterion gives teams a practical boundary task boundary for where an agent should take over and where it should not.

Generalization

Build delegation policies based on the cognitive function of an output artifact, not just on task completion cost or token cost.

writing to formulate your own thoughts (which cannot be automated) and writing to report status or summarize (which should be automated)
Open source video
Practicelow noveltymoderate evidence

In a fast-changing AI product environment, empirical/prolific shipping beats academic/theoretical planning because real user behavior and telemetry are more reliable than long reasoning documents.

Why it matters

This directly affects how engineering and product resources are allocated: invest in short experiment loops, user instrumentation, and fast release mechanisms rather than heavyweight specifications and internal debate.

Generalization

Use the shortest path to a real user interaction as the primary evaluation mechanism when the underlying substrate is changing quickly.

Empirical testing beats academic planning.
Open source video
Rapid prototyping with users accelerates product-market fit.
Open source video
Practicelow noveltymoderate evidence

Teams building AI products should negotiate AI into their own daily operating cadence—not as an occasional demo but as the infrastructure they depend on—because internal dependency creates the highest-fidelity feedback loop.

Why it matters

Dogfooding at this level surfaces reliability, context, abstraction, and workflow failures that standalone evals cannot.

Generalization

Treat internal usage as a continuous production evaluation and give the product team no escape hatch back to manual work.

Dogfooding internal AI tools aggressively accelerates product refinement.
Open source video

Deep dives

4

Runtime architecture of persistent AI coworkers

Research question

Which architectural properties—durable memory, continuous workspace access, long-lived task ownership, shared context, or recovery semantics—are necessary and sufficient for an agent to function as a persistent coworker rather than a longer-running task agent?

Why

The summary names the era but does not define the underlying mechanism. Without a precise architectural specification, product teams cannot compare designs, build the right runtime, or evaluate whether their agent is actually a coworker.

Era 3 is persistent AI coworkers who get things done alongside you.
Open source video
Source video

Empirical derivation of the 2-to-3-month model capability horizon

Research question

How can an AI product team reliably distinguish model capabilities that will be real in the next two to three months from one-year roadmap hype, using evidence available today?

Why

The 2-to-3-month rule is stated as an operating principle, but not as a method. A deriveable horizon would prevent teams from oscillating between building for today’s models and betting on unproven future capabilities.

Building for current models fails.
Open source video
The only viable window for product planning is 2 to 3 months out.
Open source video
Source video

Human steering affordances for delegated AI execution

Research question

What plan visibility, assumption checks, goal handoff, and approval/rejection gates let one human steer one or more AI coworkers without becoming a full-time supervisor or bottleneck?

Why

Rowing-versus-steering redraws the human role, but the summary does not say which control-plane artifacts make steering tractable. These affordances determine whether delegation scales beyond one-off tasks.

Knowledge work is shifting from execution to steering.
Open source video
AI handles tactical execution while humans define core hypotheses and goals.
Open source video
Source video

Mechanistic boundary between writing-as-thinking and writing-as-reporting

Research question

Can a system infer at request time whether a given writing task is cognitive (forming the author’s own thoughts) or communicative (reporting status or summarizing), and what are the downstream quality effects of routing each type differently?

Why

Without such a classification, delegation policies will automate the highest-value cognitive portion of knowledge work in the name of efficiency, silently eroding the thinking that originally gave the work its value.

writing to formulate your own thoughts (which cannot be automated) and writing to report status or summarize (which should be automated)
Open source video
Source video

Article ideas

4

The Third Era Isn’t Chat Plus Memory

Persistent AI coworkers are a new product category with architectural requirements—durable context, shared workspace access, and recovery from interrupted work—that cannot be met by adding memory to an episodic agent.

Angle

Architecture critique of current agent frameworks; argues for designing a long-lived coworker runtime instead of a request/response service.

Source video

Your 12-Month AI Roadmap Is Fiction: Build a 90-Day One

AI product teams should plan around the frontier model of two or three months from now rather than today’s model or a one-year forecast, because both extreme horizons predictably produce wasted engineering work.

Angle

Planning and build strategy for AI application teams; makes the horizon concrete and proposes capability-slotted roadmaps.

Source video

Automate the Status Reports, Never the Drafts

Knowledge workers should delegate reporting and summarizing to AI but keep drafting and outlining human, because the writing-as-thinking act is the source of insight and cannot be automated without destroying it.

Angle

Work-practice argument against blanket AI writing delegation; gives a simple cognitive-ownership test.

Source video

Eat Your Own AI—Forever

Aggressive dogfooding of internal AI tools is not a demo exercise but a production-grade evaluation: the product team’s daily dependency on the AI exposes reliability, context, and workflow failures that offline evals cannot.

Angle

Operational manifesto for AI product teams; positions dogfooding as the highest-fidelity feedback loop.

Source video

Project ideas

4

Long-Horizon Agent Persistence Harness

beyond-evals

An agent with durable checkpoints and recovery semantics will complete interrupted multi-session knowledge-work tasks at least 50% more often than a stateless agent that is restarted from scratch on each new request.

Proof of concept

Implement the same multi-step task (research, summarize, draft status update) in two agent configurations: stateless per-request and persistence-checkpointed across sessions. Randomly interrupt both configurations at equivalent points, then resume and measure completion.

Measurement

Task completion rate, average human rework time after interruption, and number of failed recovery attempts across 50 simulated task runs.

Source video

Steerable Execution Gate Trial

gatehouse

Exposing an agent’s plan and assumptions to a human reviewer before execution will raise final artifact quality by at least 25% compared with a black-box agent execution of the same task.

Proof of concept

Build a two-arm tool using the same model/tool chain for a delegated report-generation task. In the control arm, the agent executes and emits a final artifact. In the treatment arm, the agent first emits a plan plus assumptions, receives human approval/correction, then executes. Compare artifacts.

Measurement

Blinded human-rated artifact quality and rate of human correction/rollback across 20 participants and standardized knowledge tasks.

Source video

Capability Isolation Adapter

new

When model-dependent capabilities are isolated behind an adapter interface, a frontier model upgrade causes less than half the rework of an architecture where model calls are embedded directly in business logic.

Proof of concept

Implement the same miniature AI product twice: one branch with direct model calls and one with an internal capability interface plus adapter. After both are working, swap the underlying model and record the changes required to restore equivalent behavior.

Measurement

Files changed, pull-request size, and functional regression count for each architecture across two simulated model upgrades.

Source video

Thinking vs. Reporting Writing Router

new

A routing policy that sends report-style writing to full AI automation and thinking-style drafting to AI-assisted scaffolding will match user preferences and preserve self-rated thought clarity in at least 80% of tasks.

Proof of concept

Build a small assistant that classifies each incoming writing task as reporting or thinking. Reporting tasks are auto-produced; thinking tasks get outlines and partial scaffolds but not a final draft. Have participants do the same tasks with a fully automating assistant as a control.

Measurement

User preference between routing policies, self-reported clarity of thinking after each session, and agreement rate between the classifier and human labels.

Source video

Architectural implications

4

Persistent AI coworkers extend agents beyond a single task lifecycle.

Before

Agent system is instantiated per request or task; state resets when the task ends.

After

AI coworker maintains durable state, shared workspace, and continuing ownership of work across sessions.

Consequence

Runtime infrastructure must add durable memory, long-running task control, audit trails, and recovery mechanisms for interrupted work.

Source video

The 2-3 month capability horizon makes model-specific assumptions a temporary layer.

Before

Architecture treats the current model's capabilities as a stable platform and builds long-term features around it.

After

Architecture assumes capability shifts in a few months and keeps model-dependent interfaces isolated.

Consequence

Feature capabilities should be gated or adapter-based so assumptions at one capability level can be swapped without rewriting the product.

Source video

Rowing versus steering moves humans from executing work to defining goals and reviewing outcomes.

Before

Human does step-by-step execution; AI produces drafts or suggestions for each step.

After

AI executes multi-step work while human supplies hypotheses, strategic direction, and final judgment.

Consequence

Agent systems need explicit steering affordances: plan visibility, assumption checks, and human approval/rejection gates over major artifacts.

Source video

Writing-as-thinking versus writing-as-reporting makes a cognitive distinction between generated artifacts.

Before

All written output is considered equally delegable to AI.

After

Reports and summaries are automatically generated; thinking documents remain human-authored, with AI used only as support.

Consequence

Workflow tooling should label and route artifacts differently depending on whether they are meant to externalize cognition or to communicate status.

Source video

Tradeoffs and failure modes

3

Building for current model capabilities

Benefit

Immediate constraints are known and implementations can be validated against today's models.

Cost or risk

Products can become obsolete before they reach users as model capabilities shift.

Building for current models fails.
Open source video
Source video

Building for a one-year model roadmap

Benefit

Appears future-proof and ambitious.

Cost or risk

Projections are too speculative and engineering effort may be wasted on capabilities that do not arrive as assumed.

Building for where models will be in a year fails.
Open source video
Source video

Delegating writing-as-thinking to AI

Benefit

Maximizes automation and reduces time spent drafting.

Cost or risk

Removes the cognitive act that creates insight; the source explicitly says this kind of writing cannot be automated.

writing to formulate your own thoughts (which cannot be automated)
Open source video
Source video

Open questions

4

What architectural properties make an AI agent a 'persistent coworker' rather than an agent that happens to run longer? Is it durable memory, continuous context, shared workspace access, task ownership, or something else?

Why unresolved

The summary identifies the era but does not define the mechanism by which persistence is achieved.

Research direction

Build a taxonomy and evaluation set for long-lived agent state, state recovery, and cross-session task continuity.

Source video

How can a builder know which model capabilities will be real in the next 2 to 3 months and which are roadmap hype?

Why unresolved

The 2-3 month horizon is given as an operational rule, but the source does not specify how to derive or verify the horizon.

Research direction

Instrument early model previews with actual product tasks and track how product adoption changes as capability boundaries move.

Source video

What interaction and control abstractions let one human steer multiple AI coworkers without becoming the bottleneck?

Why unresolved

Rowing/steering describes the split but does not solve the scaling problem of human review and direction.

Research direction

Measure human supervision cost per delegated task and experiment with exception-only escalation and structured goal handoffs.

Source video

Can a system reliably decide whether a given writing task is thinking or reporting, or does that require the writer's own context?

Why unresolved

The summary gives a clear principle but gives no mechanism for automating the classification.

Research direction

Develop a delegation-policy model that infers cognitive purpose from task context and evaluates downstream quality effects.

Source video

Key claims

6
predictionVerification needed

AI products evolve from chat interfaces through task agents to persistent AI coworkers.

Evidence

AI products are evolving through three distinct phases: chat interfaces (era 1), task-oriented agents (era 2), and persistent AI coworkers (era 3).

Question

What observable product attributes define the boundary between an agent and a persistent coworker?

Source video
causalVerification needed

Building for current model capabilities fails and building one year out also fails.

Evidence

The only viable window for product planning is 2 to 3 months out.

Question

Can product teams using longer or shorter horizons be compared on shipped impact or rework rate?

Source video
comparativeVerification needed

Empirical testing and rapid prototyping beat academic planning in AI product development.

Evidence

Empirical testing beats academic planning.

Question

What metrics separate empirical from theoretical approaches in practice, and do they generalize outside product management?

Source video
predictionVerification needed

Knowledge work will shift from human execution to human steering, with AI doing tactical execution.

Evidence

AI handles tactical execution while humans define core hypotheses and goals.

Question

Which knowledge-work roles show measurable shifts in time spent executing versus setting direction?

Source video
opinionVerification not requested

Writing to formulate thoughts cannot be automated, while reporting and summarizing should be automated.

Evidence

writing to formulate your own thoughts (which cannot be automated) and writing to report status or summarize (which should be automated)

Source video
causalVerification needed

Aggressive internal dogfooding accelerates product refinement.

Evidence

Dogfooding internal AI tools aggressively accelerates product refinement.

Question

How does defect discovery rate or iteration speed change when internal usage is the primary evaluator?

Source video

Connections

4