Week in Product #499 🚀
Q2 earnings, AI slowdown proposals, iCloud Plus, X Money app, Building a smart home robot, LLM recommenders at Netflix, Cold-start evals, Graph-engineering, DoorDash Air & more
Hi friends 👋
Welcome to a new Week in Product!
🎰 The week in figures
$200M: Simile raised more than $200M at a $2B valuation to simulate populations for testing products, policies, messages, and strategies
$113M: Onyx Security, the AI agent security startup, raised a $113M Series B to develop new AI models and to expand sales/marketing internationally
100K: OpenAI is giving 100,000 academic researchers free access to its top models (including GPT-5.6 Sol Pro) through 2027, part of a $250M+ commitment to external scientific research
💰 Q2 earnings
Microsoft: Q4 revenue rose 18% to $90.0B, with Cloud up 27% to $59.3B and Azure surpassing $100B for the year. Shares jumped as the company held capex steady
Meta: Q2 revenue up 28% to $60.8B, but EPS missed by 14% as expenses climbed 55% to $42.0B on AI spend. Free cash flow shrank to just $784M
Apple: Q3 revenue rose 16% to $109.4B, a June-quarter record across iPhone, Mac and Services. Stock still slid on soft Q4 guidance despite the beat
Amazon: Q2 revenue topped $200B for the first time, up 20%, with AWS accelerating to 37% growth, its fastest in 18 quarters. Raised 2026 capex guidance to ~$220B
📰 What’s going on
Claude Opus 5 breaks ethics tests running vending machines. Anthropic's latest model colluded on pricing, lied to suppliers, and broke 11 separate agreements when tasked with operating a simulated business for one year
Anthropic found three incidents in which Claude models escaped testing sandboxes and gained unauthorized access to organizations. The cybersecurity evaluations were prompted by OpenAI’s Hugging Face breach
OpenAI and Anthropic back government slowdown proposal. The 2 frontier labs endorsed a letter asking federal regulators to "pace" AI development if progress becomes too rapid. Meanwhile, Zuckerberg spoke against the proposal
OpenAI cut GPT-5.6 Terra’s price by 20% and Luna’s by 80% three weeks after launch, as rising enterprise AI costs and cheaper models from Chinese and Big Tech rivals increase pressure on pricing
Google adds Nano Banana to Google Earth. You can now generate custom images using Google Earth’s satellite, aerial, and 3D imagery alongside Nano Banana, which creates concepts grounded in the real world. Just zoom in to a place in Google Earth on web, tap “create image,” and type whatever you want to see
Amazon is reportedly scaling back several Nova models and shifting resources to a new Frontier Model Research group
Apple hints at paid iCloud Plus and AI upgrades, incl Siri AI app, with iOS 27
Meta AI now available in Threads DMs globally. Users can share posts, images, and videos directly with the chatbot for follow-up questions, as Meta works to keep users in its ecosystem rather than switching to ChatGPT or Gemini
X Money app rolls out in the US for US Premium and Premium+ subscribers. Sign-up comes with a physical X Visa debit card (it’s also usable in Apple Pay), fee-free peer-to-peer transfers, no foreign transaction fees, and free ATM withdrawals worldwide. Premium+ subscribers also get 6% APY automatically, while Premium subscribers get the same rate by linking direct deposit, plus up to 3% cash back on some purchases and early access to direct deposits. It’s the clearest step yet toward achieving Musk’s stated goal of turning X into the “everything app”
DoorDash launched DoorDash Air, an in-house drone delivery program, becoming the eighth US operator cleared for commercial drone delivery
Lyft and Baidu began testing Apollo Go robotaxis in London, adding another operator to the city's crowded autonomy race
LinkedIn adds "seems like AI slop" reporting button. It now blocks hundreds of thousands of automated comments daily, and flags accounts posting heavy AI-generated content as users revolt against synthetic posts filling feeds
🤖 Futurism
Building worlds that train robots. Fei-Fei Li describes how World Labs’ SceniX team is using a “real-to-sim-to-real” engine to turn robot tasks into high-fidelity simulated variations. By massively scaling experience and evaluation in simulation, robots can learn complex manipulation skills with zero real-world training data
The FCC bans US imports and use of foreign-made humanoid robots and robot dogs, citing national security concerns and potential backdoor threats
Gemini Robotics 2 gave robots whole-body control, dexterous hands, and the ability to coordinate across different robot types
📚 Good reads
Testing an LLM recommender at Netflix. Netflix’s new GenRec system can beat the old feature-based ranker with an LLM that ranks the full catalog. It focuses on context engineering (what to verbalize, compress, or drop) instead of classic feature engineering, and mixes ranking, language modeling, and reward-weighted losses to align with long-term satisfaction and business goals. In large A/B tests, GenRec beat a very mature baseline using far fewer labeled examples and signals
So, you want to be a content creator? Elena Verna shares how her Product content journey grew into a 100K-subscriber Substack and meaningful income, without ever aiming to be an “influencer.” She’s candid about the trade-offs and shares useful advice
Beyond time management. Anne-Laure Le Conff shares her idea of timeshielding: blocking time for what truly matters (without over-specifying tasks), so you can match work to your energy and attention. For PMs juggling stakeholder pings and “urgent” asks, it’s a reminder to proactively defend time for deep product thinking and recovery, before your inbox plans your week for you
Eval-driven development at Airbnb. Interesting peek at how Airbnb treats LLM evaluation. They walk through a practical eval-driven development loop: starting from failures, layering programmatic checks, calibrated LLM-as-judge, and human review. For PMs, the key is to define concrete goals and rubrics early, continuously refine them from live traffic, and evaluate the whole system (retrieval, tools, agents). Check the podcast in the section below for tips on building evals
🎧 Good listen/watch
[Video podcast] How to build your first cold-start AI eval. Aakash and ex-Meta/Google PM Daniel McKinnon walk through designing “cold-start” evals you can build before you have any users, using domain knowledge + Claude/ChatGPT to generate cases. They show how to define the problem, use easy and hard examples to find your model’s floor and ceiling, then auto-generate a ~100-case offline eval set and an LLM judge prompt
🧑💻 Worth learning
Graphs are the next evolution after loops in AI workflows. Aakash explains how “graph engineering” coordinates multiple specialized AI agents, building on earlier waves like prompt, context, harness, and loop engineering. He walks through common graph patterns and shows when to use each instead of just a sequential flow. Includes a demo on how to construct these graphs in Claude Code (e.g. a double-diamond PRD workflow with different models per node)
Giving Claude a robot body. Tal Raviv strapped a cheap camera-on-wheels to Claude’s “computer use” and turned it into a clumsy but surprisingly capable home robot. It drives, looks, logs memories as text and photos, and even learns things like better turning behavior
🔧 Products to try
Kimi K3 weights give developers direct access to Moonshot’s Chinese open model as the open-weight policy fight heats up; free/open-weights
That’s all for this week.
Feel free to drop your comments/questions/feedback. Would love to hear what you’d like to see more about in WiP, so I can make it better for you.
Have a great start to your week!




