Silicon Valley Girl · Published 2026-09-08

Superhuman CEO: How to Position Yourself Now Before the Next AI Phase (2026–2027)

Open on YouTube ↗

Summary

Overview

  • Speaker: Shishir Mehrotra
  • Channel: Silicon Valley Girl
  • Main topic: AI's impact on employment, workflows, management skills, and career positioning in the AI era.
  • Purpose: To educate professionals and founders on how to position their careers and build AI workflows before the next wave of automation changes the job market. Shishir Mehrotra, CEO of Superhuman and former YouTube product leader, discusses how AI is transforming professional roles by shifting skills upward. Rather than replacing jobs, AI pushes the execution layer up to problem and solution finding. He introduces frameworks like PSHE and eigenquestions to help professionals stand out, avoid the recruiting funnel, and build systems around AI rather than just using it as a chat tool.

Topic Map

The Shift in Job Roles and Management

  • Explanation: AI is turning professionals into managers of workflows, tools, and agents much earlier in their careers.
  • Key claims:
    • AI pushes the career ladder up.
    • The original skill doesn't go away, it shifts.
    • The job you do today and the job you get promoted for aren't the same anymore.
  • Examples:
    • Transitioning from writing individual pieces of copy to managing AI writing agents and editorial direction.
  • Terminology:
    • AI superhighway
    • workflows
    • agents
  • Why it matters: Understanding this shift helps professionals acquire the judgment and management skills needed to stay relevant.

How to Stand Out and Avoid the Recruiting Funnel

  • Explanation: The best career transitions happen through interacting on projects and being interesting in the world, not through standard application folders.
  • Key claims:
    • Avoid the recruiting folder entirely.
    • The most interesting jobs happen through interaction.
    • Be interesting in the world by starting projects and writing things.
  • Examples:
    • Shishir joining Spotify after writing a paper called Formats as the Bundling that caught Daniel Ek's attention.
  • Terminology:
    • recruiting folder
    • formats as the bundling
  • Why it matters: Traditional job hunting puts candidates into massive competitive pools where standing out is difficult.

The PSHE Framework for Career Ladders

  • Explanation: A framework mapping out the ownership of Problem, Solution, How, and Execution across a career.
  • Key claims:
    • Early in careers, people mostly move on the scope axis at the execution level.
    • In the middle of careers, the curve reverses into the trough of disillusionment.
    • Senior roles focus heavily on problem definition and framing.
  • Examples:
    • Engineers and marketers moving from execution tasks to problem definition.
  • Terminology:
    • PSHE
    • trough of disillusionment
    • scope axis
  • Why it matters: Provides a clear mental model for how careers progress and where value is added.

Eigenquestions: The Art of Framing Problems

  • Explanation: Eigenquestions are the most discriminating questions in a set; when answered, they answer most of the other questions.
  • Key claims:
    • The hard part is not finding the right answer, it is asking the right question.
    • Eigenquestions can be practiced in low-stakes environments.
    • Children often ask great eigenquestions because they cut to the core safety or viability constraints.
  • Examples:
    • Asking whether a teleportation device is safe for humans before worrying about its capex or opex.
  • Terminology:
    • eigenquestions
    • eigenvectors
    • framing problems
  • Why it matters: Mastering eigenquestions elevates problem-solving and strategic thinking.

Building Systems Around AI

  • Explanation: The competitive advantage is not just using AI tools, but knowing how to connect data, memory, and state to build persistent agents.
  • Key claims:
    • Real AI agents need current context, reliable retrieval, memory, and state.
    • MongoDB Atlas serves as a data layer for AI apps and agents without needing completely new stacks.
    • Superhuman Go allows building AI agents that work right where you work.
  • Examples:
    • Knowledge Checker agent reviewing text against internal source of truth documents in Google Docs and Gmail.
  • Terminology:
    • MongoDB Atlas
    • Superhuman Go
    • vector search
    • retrieval
  • Why it matters: Moving beyond chat interfaces to integrated AI agents transforms personal and team productivity.

Key Points

AI turns individual contributors into managers earlier

  • Explanation: Because AI handles execution tasks rapidly, humans must focus on setting context, managing workflows, and exercising judgment.
  • Evidence: Grammarly generating over 100 billion LLM queries a week at thousands per user daily.
  • Practical implication: Professionals must practice judgment, editing, and problem framing rather than just raw execution.

Practice skills in low-stakes environments

  • Explanation: Complex skills like judgment and problem framing are best learned through low-stakes repetition before facing high-stakes situations.
  • Evidence: Learning musical instruments or sports through practice rather than immediate recitals.
  • Practical implication: Do side projects and low-stakes experiments with AI tools to build competence.

Assist vs Chat vs Do metaphors

  • Explanation: AI interaction models range from chat (conversational agents) to do (task lists) to assist (agents that work alongside you proactively).
  • Evidence: Superhuman Go providing inline assistance and triggers without requiring users to prompt manually.
  • Practical implication: The most valuable AI tools are assistants that understand context and work proactively where you work.

Frameworks, Models & Processes

PSHE (Problem, Solution, How, Execution)

  • How it works: Maps career progression from executing given solutions to defining problems and framing the strategic landscape.
  • Components:
    • Problem
    • Solution
    • How
    • Execution
  • When to use: When evaluating career growth, promotions, and how responsibilities shift under automation.

Eigenquestions

  • How it works: Identifies the single most discriminating question that resolves downstream questions when answered.
  • Components:
    • Identifying the set of questions
    • Testing for discriminant power
    • Focusing on framing over answering
  • When to use: During complex problem-solving, product design, and strategic decision-making.

Examples & Case Studies

Shishir wrote a paper called Formats as the Bundling which was passed around Daniel Ek.

  • Illustrates: How to stand out and avoid the recruiting folder by putting ideas into the world.
  • Lesson: Be interesting publicly to create organic career opportunities.

Google infrastructure leader mapped her team on scope and PSHE axes.

  • Illustrates: The trough of disillusionment where professionals experience career identity crises as scope expands.
  • Lesson: Understand that hitting the middle management transition requires shifting from execution to problem framing.

Actionable Takeaways

  • Immediate:
    • Practice asking eigenquestions in low-stakes settings.
    • Build side projects to showcase your skills outside the traditional recruiting funnel.
  • Strategic:
    • Shift your focus from execution to judgment, curation, and problem framing.
    • Integrate AI assistants directly into your daily workflows rather than treating them as separate chat tools.
  • Questions to investigate:
    • How can I shift my current job responsibilities from execution to problem definition?
    • What data sources and context do my AI agents need to be genuinely autonomous and reliable?

Claims Worth Verifying

  • Grammarly handles over 100 billion LLM queries a week. (statistical)
  • Grammarly acquired Superhuman in July 2025 and rebranded in late 2025. (corporate history)

Notable Quotes

"You learn to be a manager. I spent years giving people the opposite advice." (at 0:00) "The job you're doing today and the job you get promoted for aren't the same anymore." (at 0:27) "The hard part is not about finding the right answer, it's about asking the right question." (at 27:06)

Compressed Summary

  • AI shifts career skills upward from execution to problem framing and judgment.
  • The PSHE framework explains how professionals transition through execution, how, solution, and problem ownership.
  • Eigenquestions are the most discriminating questions that answer all downstream questions.
  • Avoid the recruiting folder by putting work and ideas publicly into the world.
  • MongoDB Atlas provides the data layer for AI agents requiring context, memory, and vector search.
  • Keywords: artificial intelligence, career progression, eigenquestions, management, workflows
  • Core insight: AI accelerates execution and pushes human value upward into problem framing, judgment, and system management.

Core insights

4
Architecturehigh noveltymoderate evidence

Agent 'assist' mode is fundamentally different from chat or task-list execution: it is event-triggered, runs inline within the user's existing tools, and requires no manual prompting. Building an agent as a chat application caps its value and requires an event/trigger infrastructure plus a permissions/context layer.

Why it matters

The interaction model determines control flow (request/response vs event-driven), context access, and permission boundaries. Proactive assistants that work where users already work can reach far higher utilization than chat-only agents.

Generalization

When designing any agent that operates over user documents, email, or IDEs, consider offering an assist mode driven by workspace events rather than only a conversational loop.

AI interaction models range from chat (conversational agents) to do (task lists) to assist (agents that work alongside you proactively).
Open source video
Superhuman Go providing inline assistance and triggers without requiring users to prompt manually.
Open source video
Architecturemedium noveltymoderate evidence

Real agent memory/state requirements can be satisfied by extending an existing operational database with vector search, instead of introducing a separate AI stack. This gives agents durable context, current state, and reliable retrieval with transactional consistency.

Why it matters

Choosing where to store state and embeddings determines whether retrieval and state updates can be atomic, and whether you need to move data between systems. A single operational data layer simplifies agents and makes them more reliable in production.

Generalization

Before adopting a purpose-built vector store or agent-memory framework, evaluate whether an existing OLTP database with vector search can satisfy the retrieval and state requirements.

Real AI agents need current context, reliable retrieval, memory, and state.
Open source video
MongoDB Atlas serves as a data layer for AI apps and agents without needing completely new stacks.
Open source video
Mental Modelhigh noveltymoderate evidence

As AI absorbs execution tasks, the highest-leverage human or agent operation shifts upward to problem framing. The key is identifying the 'eigenquestion'—the most discriminating question whose answer resolves the largest number of downstream questions—before spending effort on solutions.

Why it matters

For agentic systems, this implies that reliability and usefulness depend less on answer generation and more on planning stages that elicit constraints and frame the problem correctly. Tools that help users ask the right questions may matter more than better generators.

Generalization

When decomposing a complex agent task, first enumerate candidate questions and test which answer would collapse the planning tree most, then focus compute on that question.

Eigenquestions are the most discriminating questions in a set; when answered, they answer most of the other questions.
Open source video
The hard part is not finding the right answer, it is asking the right question.
Open source video
Empirical Resultmedium noveltyweak evidence

At 'assist' scale, an embedded AI product can generate thousands of LLM calls per user per day—the summary cites Grammarly at over 100 billion LLM queries per week. This makes per-query cost, latency, caching, and retrieval first-order engineering constraints rather than afterthoughts.

Why it matters

If proactive agents reach such high per-user invocation rates, architectures must be optimized for marginal cost and latency; treating every call as an independent generation is not viable.

Generalization

Estimate an agent product's load not from user sessions but from the number of triggers/actions per user per day, and design for high-rate, low-latency operation.

Grammarly generating over 100 billion LLM queries a week at thousands per user daily.
Open source video

Deep dives

4

Operationalizing eigenquestions for agentic planning

Research question

Can an agent discover the eigenquestion for an underspecified task by ranking clarifying questions with expected information gain, and does asking that question materially reduce planning cost and improve plan quality?

Why

As AI takes over execution, remaining human/agent leverage lies in problem framing. A computational definition of eigenquestions would let orchestrators stop wasting inference and user attention on low-discriminating questions and instead ask the one question that collapses the solution space.

Eigenquestions are the most discriminating questions in a set; when answered, they answer most of the other questions.
Open source video
The hard part is not finding the right answer, it is asking the right question.
Open source video
Source video

Permission and context architecture for event-driven assist agents

Research question

What scoped, origin-aware event-selection and permission model lets an assist agent operate inline in email/docs without manual prompting while preventing unwanted or stale-context actions?

Why

Assist-mode agents unlock value by acting where users already work, but removing the manual prompt shifts the risk to context drift and unauthorized operations. Without a provenance and permission architecture, proactive actions will not be safe enough to deploy on private data.

AI interaction models range from chat (conversational agents) to do (task lists) to assist (agents that work alongside you proactively).
Open source video
Superhuman Go providing inline assistance and triggers without requiring users to prompt manually.
Open source video
Source video

Cost and latency engineering at assist scale

Research question

What triggering, caching, and retrieval policies can keep an assist-mode agent's cost/latency viable at thousands of LLM calls per user per day while maintaining user attention and suggestion relevance?

Why

Embedded AI products can generate far more calls per user than chat tools, so marginal cost and latency become first-order product constraints; the architecture must be designed for high-rate, low-latency operation rather than treating every call as an independent generation.

Grammarly generating over 100 billion LLM queries a week at thousands per user daily.
Open source video
Source video

Empirical comparison of integrated vector search vs purpose-built vector DB for agent memory

Research question

Can an operational OLTP database with integrated vector search match dedicated vector databases' recall and latency for agent hybrid queries (structured filters + similarity) under concurrent writes and high per-user invocation?

Why

Architectural preference for one data store depends on demonstrable index and tail-latency performance. If a single operational database is enough, agents become simpler and more reliable; if not, the extra dedicated AI stack is justified.

Real AI agents need current context, reliable retrieval, memory, and state.
Open source video
MongoDB Atlas serves as a data layer for AI apps and agents without needing completely new stacks.
Open source video
Source video

Article ideas

4

Stop Building Chat Agents: The Next Step Is Assist-Mode Architecture

Agents that can only be invoked through a separate chat UI miss the context and friction problem; the winning pattern is an event-triggered assistant embedded in the tools where work already happens, so product and architecture decisions should center on event buses and scoped permissions, not prompt polish.

Angle

Technical argument for shifting agent products from request/response to event-driven inline assistance

Source video

Agents Don't Need Better Generators—They Need Eigenquestions

As AI absorbs execution, the remaining bottleneck is upstream problem framing; agent systems that spend inference budget on asking the right discriminating question before decomposing a task will outperform systems that invest the same budget generating answers from underspecified prompts.

Angle

A contrarian architecture argument for adding a question-discovery stage to orchestrators

Source video

Grammarly-Scale Means Cost Is a Product Decision

Once an assist-mode product reaches thousands of LLM calls per user per day, per-query cost and latency are not infrastructure details but determine what triggers are sent and which suggestions are worth surfacing; durable agents must be engineered like high-QPS systems with aggressive caching and cheap retrieval.

Angle

Scale analysis of embedded assistants, using Grammarly's stated query volume as upper bound

Source video

One Database for Agent Memory Is the Pragmatic Architecture

For assistants that must reconcile fast-changing state with semantic retrieval, keeping embeddings next to operational data in an existing database with vector search beats bolting on a separate AI memory stack—until a benchmark shows a purpose-built index actually wins on recall or tail latency.

Angle

Architecture tradeoff argument against premature infrastructure for agent memory

Source video

Project ideas

4

Question Value Probe: Ranking Clarifying Questions as Eigenquestions

beyond-evals

A question-ranking model trained to maximize expected reduction in solution-space branching will select the expert-judged eigenquestion at least 70% more often than a baseline that asks for generic clarification.

Proof of concept

On an existing set of underspecified coding/planning tasks, enumerate candidate questions, generate possible solution plans for each candidate, compute average pairwise plan edit distance as a proxy for solution-space reduction, and compare predicted top question to human-annotated eigenquestion.

Measurement

Top-1 match with expert ranking; improvement in downstream task success when the agent asks the top predicted question before planning, versus baseline question order.

Source video

Tightrope: Permissioned Assist Agent in a Live Editor

gatehouse

An event-triggered agent that only sees the current buffer/document and surfaces one inline suggestion will reduce the number of user actions required to complete a document-editing workflow by at least 40% compared with a chat agent that requires context pasting, with no increase in user-corrected edits.

Proof of concept

Build a minimal VS Code extension or browser email client plugin that fires on save/typing/compose events within a single document scope, runs a knowledge-check prompt, and shows accept/reject inline diffs; run the same task via chat-only baseline.

Measurement

Number of user actions and prompts, accepted vs user-corrected suggestion rate, user-reported confidence in 10 participants.

Source video

One-Store Agent: Unified Operational + Vector Store Benchmark

beyond-evals

For a realistic email/document agent workload, a single operational DB with vector search can match a purpose-built vector DB within 5 percentage points of recall@10 and within 2x p95 latency, with atomic state updates and a simpler ops footprint.

Proof of concept

Take a snapshot of documents/emails and generate embeddings; replay hybrid queries where filters restrict on source/status and distance scores rank semantically against a purpose-built vector database running the same embeddings; concurrently write state updates to both stores.

Measurement

recall@10, p95/p99 latency, throughput under simultaneous writes, operational complexity in number of services and sync jobs.

Source video

Trigger Budget: Utility-Gated Proactive Assist

new

If an assist agent triggers only on events for which an inexpensive local model predicts high suggestion acceptance probability, it can match all-event triggering in accepted suggestions per user session while reducing daily LLM calls and cost by more than 50%.

Proof of concept

Extend the Tightrope assist-mode prototype with a trigger gating classifier using local features such as event type, text diff size, recency, and document metadata; compare all-event triggers vs gated triggers over the same user workflow tasks.

Measurement

Accepted suggestions per session, dismissed/undo rate, LLM calls per user per day, user opt-out rate.

Source video

Architectural implications

3

Agents that live in a separate chat UI cannot see the user's current document/mail context and require manual copy-paste or context provision.

Before

Users open a chat window and explicitly provide context every time they want agent assistance.

After

Agents run inline in editors and clients, subscribe to workspace events, and receive context automatically.

Consequence

Architectures must shift from purely synchronous request/response to event buses, webhooks, or pub/sub triggers, with fine-grained scoped permissions.

Source video

Agent state is often prototyped as ephemeral context or in a separate vector store, leading to consistency issues between retrieved knowledge and operational data.

Before

Memory/state is kept in LLM context or dedicated vector DB; transactional data is synchronized separately.

After

Use a single operational database with integrated vector search and ACID semantics to store state, embeddings, and app data together.

Consequence

Retrieval results and state updates can be atomic, simplifying rollback and consistency for agents that read and write user data.

Source video

The bind on agent work is upstream: defining the problem and identifying the key question, not generating the final artifact.

Before

Agent workflows are one-shot: user prompt -> generated answer.

After

Agent workflows include an explicit problem-framing step that asks the eigenquestion before decomposing into execution tasks.

Consequence

Orchestrators should route through a 'question discovery' module that can ask the user for the one constraint that most reduces solution space.

Source video

Tradeoffs and failure modes

2

Proactive assist vs explicit user-initiated chat

Benefit

Removes prompting friction and acts where work happens, enabling very high usage.

Cost or risk

Without a manual prompt, the agent may act on stale context, perform unwanted operations, or require broad data permissions that create privacy/security exposure.

Superhuman Go providing inline assistance and triggers without requiring users to prompt manually.
Open source video
Source video

Using an existing operational database as agent memory/vector store

Benefit

Avoids a second data stack and keeps state and retrieval consistent.

Cost or risk

General-purpose databases may have weaker vector-search scalability or specialized ANN index tuning compared with purpose-built vector databases.

MongoDB Atlas serves as a data layer for AI apps and agents without needing completely new stacks.
Open source video
Source video

Open questions

4

How can an agent systematically generate and rank candidate eigenquestions for an underspecified user task?

Why unresolved

The summary presents eigenquestions as an expert skill, but provides no computational method for discovering them.

Research direction

Formulate question selection as expected information gain or decision-theoretic value of information and test it against planning benchmarks.

Source video

What is the appropriate permission and privacy architecture for event-driven 'assist' agents that operate inside email and documents without manual prompting?

Why unresolved

The product example (Knowledge Checker in Google Docs/Gmail) implies deep access, but governance/consent is not discussed.

Research direction

Design scoped, origin-aware permissions with user auditable inline actions, and compare usability/safety trade-offs.

Source video

What latency and cost budgets are acceptable for an assistant that triggers on every relevant workspace event without exhausting user attention?

Why unresolved

Grammarly's huge query volume hints at scale but not at throttling or user-notification policies.

Research direction

Benchmark triggering policies and measure user engagement/annoyance as a function of assist frequency.

Source video

Can an existing OLTP database with vector search match purpose-built vector stores for hybrid queries (structured filter + semantic similarity) at the throughput required by assist agents?

Why unresolved

The comparison claim is not supported by any benchmark data.

Research direction

Run comparative evaluations on real agent workloads measuring p95 latency, recall@k, and concurrent update performance.

Source video

Key claims

6
factualVerification needed

Grammarly generates over 100 billion LLM queries per week, with thousands per user daily.

Evidence

Grammarly generating over 100 billion LLM queries a week at thousands per user daily.

Question

Can this query volume be independently audited, and over what measurement window and user base?

Source video
opinionVerification needed

Real AI agents need current context, reliable retrieval, memory, and state.

Evidence

Real AI agents need current context, reliable retrieval, memory, and state.

Question

Are there classes of agents that can operate acceptably with stateless, retrieval-free designs?

Source video
comparativeVerification needed

MongoDB Atlas serves as a data layer for AI apps and agents without needing completely new stacks.

Evidence

MongoDB Atlas serves as a data layer for AI apps and agents without needing completely new stacks.

Question

What limitations exist compared with purpose-built vector databases or agent-memory frameworks in scale and index quality?

Source video
predictionVerification needed

The most valuable AI tools are assistants that understand context and work proactively where you work.

Evidence

The most valuable AI tools are assistants that understand context and work proactively where you work.

Question

Does proactive, context-aware assistance produce higher long-term user engagement than chat-based tools in controlled studies?

Source video
causalVerification needed

Eigenquestions are the most discriminating questions in a set; when answered, they answer most of the other questions.

Evidence

Eigenquestions are the most discriminating questions in a set; when answered, they answer most of the other questions.

Question

Can the concept be operationalized into a metric (e.g., information gain) and validated on real decision-making tasks?

Source video
predictionVerification needed

AI pushes career value upward from execution to problem and solution finding.

Evidence

Rather than replacing jobs, AI pushes the execution layer up to problem and solution finding.

Question

Will this hold in domains where AI also improves at problem decomposition and framing?

Source video

Connections

4