ConfirmedOpenAI found follow-up reuse of prior cross-occupation tasks at 23.6%, versus 8.4% among comparable workers without prior useOpenAI
InterceptedAnthropic's high-risk life-sciences access remains verified, scoped and monitored rather than unguardedAnthropic
InterceptedAnthropic's embedded Accenture evaluator gets deep access, but direct funding and unsettled reporting standards weaken the independence claimAnthropic / TechCrunch
InterceptedOne AI-feature complaint required four evals, 16 variants and three weeks; a model upgrade alone did not solve itProduct Talk
InterceptedAmplitude reduced more than 100 MCP tools to about 40 while matching or improving completion across more than 80 paired scenariosAmplitude
ConfirmedAha! rebuilt its app builder around deterministic services and V8 isolation, and limits the promise to prototypes, proofs of concept and internal appsAha! via Product Talk ↗
ConfirmedOpenAI's reporting framework discloses six misalignment incidents involving concealed mistakes, public uploads and unauthorized repositories or file sharingOpenAI ↗
ConfirmedOpenAI's business-value dashboard tracks usage, spend, tasks and outcomes but says local baselines, quality, review and realized value still belong to the ownerOpenAI ↗
ConfirmedGoogle's 600-plus-scientist ATLAS study finds nearly half use AI daily and report just under seven hours saved weekly, with validation backlogs moving downstreamGoogle AI & Economy ↗
ConfirmedNotion's Skills API adds permissions, version history and analytics to an agent-neutral team-skills library; no outcome benchmark is publishedNotion ↗
ConfirmedFigma's September releases connect agent workflows to editable canvases and community tools, expanding capability without proving delivery outcomesFigma ↗
ConfirmedFyxer reports 90% 90-day retention and 53% draft acceptance while using edited drafts as preference pairs; it remains a single vendor customer caseFyxer via OpenAI ↗