As models get cheap, surplus accrues to whoever owns the workflow — especially in a regulated, relationship-heavy vertical in SEA or India — not to whoever owns the model.
Year to Date — 2026
The Year in One Move
The moat kept moving up the stack — and every rung it left behind commoditized within weeks. Prompts gave way to skills (March), skills to harnesses (April), harnesses to procedures and loops (May), loops to encoded judgment — verifiers, evals, taste (June) — and by late June the frontier had moved to unit economics itself (token engineering). Through all of it, exactly one thesis stayed dead-center on the reader's edge and compounded: an AI-native managed-services company for SEA/India, in one regulated, relationship-heavy vertical. Conviction 72 → 83 across the year, Window Opening — held back only by the one thing the corpus has never produced: demand-side proof from the region itself.
The Arc
March — the skill layer crystallizes
Claude Code's skill system went from developer curiosity to mass-market operating layer in 30 days (Thariq 43.9k bm, Akshay 43.1k bm, GStack into native distribution). The founding read was already services-shaped: Levie's outsourcing frame — "the buyer is already purchasing an outcome" — plus the first unit-economics sketch (Ganim's $1–2K vertical skill libraries + $500/mo retainers). The AutoResearch loop proved any scoreable function can be automated. One Move: productize the labor that skills make executable.
April — from "use AI to code" to "use AI to run a company"
Karpathy's LLM Knowledge Bases (107k bm — the year's biggest single item) declared token throughput was migrating from code to knowledge. YC published its company-building playbook; the first revenue proofs landed (Grader's $1B in customer sales, eCom_Amin's agency inside Claude, Foaster). Meanwhile the harness commoditized in real time — GStack open-sourced, GBrain MIT-licensed. The durable edge clarified: encoded workflow knowledge plus the regional delivery relationship.
May — the operator layer gets built; the services seeds repeat
Dominant volume was pure operator-layer (CLAUDE.md at 82k stars, harness engineering, Dynamic Workflows shipping natively in Claude Code — killing "we orchestrate agents" as a moat). The edge signal ran quieter but repeated weekly: Anthropic's "Claude for Small Business," Cofounder 2, Higgsfield's agentic ad pipeline, and Isenberg's SF report — "3 billionaires are all buying SaaS companies and rebuilding them agent-first." Late May: GPT-Realtime 2.0 made voice a candidate operating surface. Services still forming, not yet actionable — the month the thesis dipped before it hardened.
June — the call gets made
The month of convergence: YC said it outright ("the biggest companies of the next decade won't be software businesses — they'll be services companies rebuilt with AI"), Dev Khare (Lightspeed India) named AI-powered services underrated in public, and two builders shipped agency-grade ad creative as a company in the same 24 hours (Rubbrband; Parker open-sourcing the agency playbook). Frontier-internal proof stacked up — Anthropic automating 95% of its own analytics, Ramp at 75%+ agent-written code. Fable 5 launched mid-month and reset the capability ceiling; loops went fully mainstream (and grifty) two weeks later. One Move: Found the SEA/India managed-services company; highest-conviction shape, done-for-you ad creative. 78/100.
July — the unit economics harden, then the model layer breaks open
The June 27 token-engineering cluster (Coinbase halving AI spend while usage climbs; JPM calling the volume shift to small/open models) is the enabling macro that separates a 70%-margin software business from a dressed-up consultancy — exactly what labor-replacement in price-sensitive markets needs. Then mid-July delivered the mechanism itself: Kimi K3, a Chinese open-weight model matching frontier on benchmarks, landed as an "AI-trade scare" (Gavin Baker: "an important inflection point… negative for Anthropic and OpenAI, net positive for everyone else") and hit #10 on OpenRouter in 48 hours — but its serving infra buckled, relocating the moat from model to inference-supply. The back half specified the founder's architecture: Etched-via-gokulr ("whoever makes the most tokens wins") and Karpathy-via-0xMortyx ("agents are distillation at scale — small model + right tools + closed loop") say the recipe is distill-small, serve-cheap, own-the-orchestration. Late July then turned the model-layer story into a live business-model fight: Jon Stokes' steelman (amplified by nic_carter to 862K impressions) — if labs can't mark up metered tokens, they don't have a business and "the USG does not owe them one" — plus Satya-via-mvanhorn ("train small in-house until it matches frontier") and Mignano-via-gokulr ("the labs won't win the app layer; value shifts to the application layer"). The surplus is now argued to accrue structurally to the workflow owner, not the model — and Anthropic open-sourcing Claude enterprise usage data handed founders a free map of which paid workflows to wrap (chrispisarski: "start an AI services agency that deploys" them). Against all of it, the first substantial regional voice in the corpus arrived as a warning, not a confirmation: H4rest4u's TaniHub post-mortem (487K impressions, Bahasa) — a celebrated Indonesian agritech "disruption" thesis that was "HALU" (delusional) on the fragmented, cash-based ground. Hold the call; conviction 80 → finalized at 82 on the August 1 mark. The margin argument is now structurally settled in the founder's favor; the demand question is the only thing left — and July's regional disconfirm (TaniHub) raised its stakes rather than answering it. Demand-side proof from the region still missing.
August — (in progress)
The services thesis acquired both its operating mechanism and its first enterprise-scale cost proof: marketing agents now sit on live business data and loop through research, action, results, and improvement; OpenAI-via-@LunarResearcher paying FDE engineers up to $785K/year exposes deployment—not model access—as the scarce labor; and Uber-via-@praveenTweets more than quadrupled frontier-AI-tool users since the start of the year while cost per token fell through caching, tighter defaults, usage visibility, and open-weight routing. Production eval operations also converged independently at Airbnb and DoorDash, turning rubric design and judge calibration into a recurring services-shaped function. DeploymentInc then supplied the first named Indian operator: @vikramchopra says its team spent two years taking AI deep into CARS24 operations and now processes more than 1.3 trillion tokens a month. That strengthens the regional delivery mechanism, but not the demand case—DeploymentInc is still supplier-side evidence, with no external buyer or willingness-to-pay proof in the bookmark. The final full week widened that contradiction: Asana-via-@WesRoth says Codex compressed an engineering migration expected to take at least five years into about two weeks for roughly $12,000, while @spakhm says ground-level enterprise adoption is still mostly unused transcription and autocomplete. Cheap capability is proven; buyer pull is not. The consumer lane weakened too: @omooretweets and @kobelum both warned that X/VC behavior badly overstates mainstream AI and wellness demand. Meanwhile @edleonklinger’s personal relationship graph proved rich context can produce non-generic action, then undercut memory as a moat by calling it “surprisingly easy to build.”
src: @gregisenberg https://x.com/gregisenberg/status/2081814601851900221
src: OpenAI-via-@LunarResearcher https://x.com/LunarResearcher/status/2083317178711921133
src: @sheherenow_ https://x.com/sheherenow_/status/2082226100764369045
src: Uber-via-@praveenTweets https://x.com/praveenTweets/status/2085124500614680891
src: Airbnb-via-@undefinedKi https://x.com/undefinedKi/status/2084279627204235703
src: DoorDash-via-@undefinedKi https://x.com/undefinedKi/status/2084743877978735014
src: @vikramchopra https://x.com/vikramchopra/status/2086696023578284489
src: Asana-via-@WesRoth https://x.com/WesRoth/status/2090053670876434917
src: @spakhm https://x.com/spakhm/status/2089719066357309671
src: @omooretweets https://x.com/omooretweets/status/2090246379679719682
src: @kobelum https://x.com/kobelum/status/2090862515043328404
src: @edleonklinger https://x.com/edleonklinger/status/2090106932916883750
Thesis Tracker
Skills and harnesses keep becoming free.
Consumer AI in SEA — especially wellness, self-knowledge, and clinic delivery — is a real lane. This corpus barely feeds it.
Memory & Context
fadingMemory and context architecture is a real capability gradient and a technical moat, not a GTM wedge.
Prosumer / Creator AI
absorbedProsumer / creator tools are demand indicators for managed delivery, not standalone companies.
Scoring note: the Mar–May briefs were backfilled retrospectively on Jun 29 (each opened marks fresh rather than deltaing a live baseline), so treat cross-month deltas as directional, not precise. Jun→Jul is a live, marked series.
Next Category evolution (the mandatory out-of-aperture pick): agent-to-agent substrate (Mar) → model-streamed UI (Apr) → voice as operating surface (May) → spatial/world models (Jun, held Jul-2) → longevity / ECM-aging biology (Aug-1 mark) — rotated in on aubreydegrey's ECM-repair breakthrough (Revel/SENS spinout) after world models produced no new corpus signal. Standing call unchanged in shape: Back, not Found; the reader can't found the biology, but the SEA clinic/delivery layer eventually rejoins the wellness edge.
The Graveyard
What the year killed, and what killed it:
- Orchestration-as-moat — died May, buried June. Dynamic Workflows shipped natively; loop directories and a Google loop-engineering PDF made even loop scaffolding free. Any "we orchestrate agents" pitch is dead.
- Memory as a Found target — decayed all year. Crowded with content, converging on a technical moat, and the context-window anti-thesis never went away. Demoted to enabler.
- Model-streamed UI (April's Next Category) — no follow-through in the corpus after the Zain Shah spike. Paradigm still plausible; category didn't form.
- Voice as operating surface (May's Next Category) — partially alive (a16z's voice-as-system-of-record, June), but no application-layer follow-through and no SEA/India product. Downgraded to watch.
- Prosumer tools as companies — commoditized continuously: Stitch free, GStack open-source, FluidVoice (local, free) killing paid Wispr Flow. The wedge was never the tool.
- The predicted regional buy-and-rebuild transaction — May's carry-forward, still unfilled in July. The billionaire pattern stayed a SF pattern. Its absence is now itself evidence (see Meta-Shifts).
- Invocation-layer-as-a-company — raised Jul 6 (Carson/Gupta: firing the right skill for a non-technical operator is the real blocker), dead by month-end. It decayed into dev-content and never became a company; once K3 commoditized the model, "invocation UX" stopped being the scarce thing. Absorbed into the services wedge as a feature, not a wedge.
- Data-licensing-as-a-thesis — the Jul 20 "sell SEA data to labs" meme (viks_rum 5.6K, omoore) produced zero operator implementation and faded to noise as a standalone thesis within a week. Survives only as a UE hedge inside a consumer wedge, not a wedge itself.
Meta-Shifts
How the corpus itself — not just its subjects — changed:
- The moat migrated up the stack, monthly. Prompts → skills → harnesses → procedures → verifiers/taste → unit economics. The consistent rule: whatever layer the discourse celebrates is 4–8 weeks from being free. The layers that never commoditized: the relationship, the encoded domain judgment, the distribution.
- From capability to economics. H1 discourse asked "what can agents do"; by late June it asked "what does a token cost." Token engineering becoming a named discipline is the year's quietest structural shift — it's what makes the services thesis pencil in price-sensitive markets.
- The maximalism correction. March–May was swarm-and-harness maximalism; June brought the pushback from inside the frontier ("don't build agents for everything" — Anthropic's own; "loops are only as good as the encoded workflow knowledge"). The field corrected toward encoded judgment over swarm size.
- The corpus's own composition is the standing risk. Every monthly flagged the same blind spots, unchanged: ~90% SF AI-builder Twitter; zero demand-side or customer voice; wellness (a named edge) essentially absent; SEA/India — the HIGHEST-edge geography — missing as a primary subject. The Found thesis is built on supply-side inference. The follow-list, not the world, is the binding constraint on this whole time-series.
- Supply-side hype vs. regional pull. The first faint disconfirming read appeared in July: the regional ad-creative competitor June predicted did not materialize. Late July sharpened it — H4rest4u's TaniHub post-mortem (487K impressions), the first substantial regional operator voice in the corpus, is a warning: a marquee Indonesian "disruption" thesis that was "HALU" on the fragmented, cash-based ground. The year's biggest open question is now explicitly testable, and the test has teeth: is there SEA/India demand pull, or only SF supply push — and does a wedge survive contact with the messy field, or only the deck?
Standing Questions Into H2
- Does a single SEA/India AI-services transaction or launch surface — the event every month since May has watched for?
- Does wellness/consumer ever appear in the corpus at investable density, or does the follow-list need surgery first?
- Does token cost keep falling on a public curve (Coinbase-style disclosures), locking in the services margin model?
- Does the world-models substrate become API-accessible — the tripwire for pre-positioning the consumer application layer?