Skip to content
All tags

#openai

28 posts

Pricing Watch | OpenAI Adds a 6x Ultrafast Speed Tier, Halves Pro 200's Usage Allowance

On 2026-09-29, OpenAI's DevDay added an Ultrafast speed tier to the GPT-6 Astra API: input/output priced at 6x Standard (short context: $60/$300 per 1M tokens) in exchange for up to 6x faster generation via the API and 8x in Codex. In the same announcement, ChatGPT Pro 200's Codex/Work usage allowance dropped from 20x Plus to 10x Plus — existing subscribers keep their old allowance until 2026-10-29, then receive a one-time $2,500 usage credit that expires 2026-12-31. The new Pro 500 plan ($500/month, 25x Plus usage) is now the only subscription tier that includes Ultrafast.

Model Card: GPT-6.1 Sol

GPT-6.1 Sol (`gpt-6.1-sol`): launched at DevDay on 2026-09-29, 1,050,000-token context window (max input 922,000, max output 128,000, same as GPT-6 Sol); standard pricing unchanged at $2.00 input / $10.00 output, but cached input is halved again from GPT-6 Sol's $0.20 to $0.10 (95% below standard input); matches GPT-6 Astra on DeepSWE v1.1 at roughly 1/5 of Astra's cost, beating GPT-6 Sol's best score by 6.4 points; beats Claude Opus 5.5 on AutomationBench at medium effort by 2.2 points at about 1/3 the cost; in the same week OpenAI shelved a flagship upgrade, GPT-6.1 Astra, over safety concerns and did not release it

Pricing Watch | OpenAI Ships GPT-6 Sol and Luna, Cutting API Prices Another 50% Off GPT-5.6's Promo Rate

OpenAI's launch post and pricing page confirm GPT-6 Sol input dropped from $4.00 to $2.00/1M tokens (-50%), output from $20.00 to $10.00 (-50%); GPT-6 Luna input dropped from $0.20 to $0.10 (-50%), output from $1.20 to $0.50 (-58%), effective 2026-09-22. The baseline for the cut is GPT-5.6's promotional pricing, not its list price — and OpenAI's own pricing page notes that GPT-5.6 Sol's promo rate is only guaranteed through 2026-11-21. That means GPT-6 Sol's new $2/$10 is the real long-term number to compare against, not the promo it's being measured against.

Model Card: GPT-6 Sol / Luna

GPT-6 Sol (`gpt-6-sol`) / GPT-6 Luna (`gpt-6-luna`) shipped 2026-09-22; 1,050,000-token context window (922,000-token max input, 128,000-token output cap — same as flagship Astra); Sol prices at $2.00 input / $10.00 output, Luna at $0.10 input / $0.50 output per 1M tokens (a further 50% cut versus GPT-5.6 promotional pricing); Sol hits 33.2% on AutomationBench at xhigh effort, beating Claude Opus 5's 26.9% at roughly 1/11 the cost; DeepSWE v1.1 puts Sol at 68.8% and Luna at 66.6%, closing in on Claude Fable 5.1's 69.9%; in OpenAI's internal simulated deployment testing, severity-3+ misalignment flags dropped from 66 to 42 versus the prior generation

Five Clouds, Five Memory APIs: How OpenAI, Anthropic, Google, AWS, and Microsoft Let Agents Remember

All five major cloud platforms shipped agent memory APIs in 2025–2026, but their design philosophies diverge sharply: OpenAI writes memory as files, Anthropic mounts memory as a directory, Google uses vectors with topic classification, AWS combines events with pluggable strategy pipelines, and Microsoft abstracts memory behind context providers. Pricing ranges from free to $0.75/1K records/month; tenant isolation spans from 'your app handles it' to IAM as a first-class citizen.

Commercial Landscape: OpenAI, Perplexity, Gemini, Claude, Grok

By 2026, the deep research commercial market has differentiated: OpenAI is comprehensive, Perplexity is fast, Gemini integrates ecosystems, Claude reasons deeply, Grok is real-time. This article compares each product's differences—not who is best, but who fits your scenario.

Multi-Agent Communication: Handoff, Delegate, Mailbox, and the Push for Protocol Standards

Agent-to-agent communication falls into three patterns: handoff (transfer control), delegate (dispatch and wait for results), and mailbox (real-time peer-to-peer messaging). Implementations vary widely, but MCP and A2A are driving protocol standardization.

Multi-Agent Context Management: The Fork vs Fresh Trade-off, History Truncation, and Result Compression

Should a sub-agent see the parent's conversation? Fork carries full history but token costs grow exponentially. Fresh saves money but lacks context. Industry consensus: default to Fresh, Fork only when needed, and always pair it with history truncation and result compression.

Multi-Agent Cost Control: How Seven Frameworks Handle the 'Soft Landing Before Hard Stop' Consensus

Parallel + nested agent spawns can burn 200K+ tokens in a single conversation turn. From Anthropic to Microsoft, the industry is converging on tiered responses: compress → downgrade → stop, rather than a binary kill switch.

The Multi-Agent Landscape: How Every Major Coding Agent Does Multi-Agent Collaboration in 2026

By 2026 nearly every mainstream coding agent supports subagents. Design philosophies split three ways: deterministic scripted orchestration (Claude Code Workflow), model-driven autonomy (Codex, Devin), and IDE command-center integration (Windsurf 2.0, VS Code). This overview maps product positioning, a capability matrix, and the design-philosophy spectrum.

Multi-Agent Orchestration Patterns: Scripted, Model-Driven, or Hybrid — How to Choose

Multi-agent orchestration splits into three camps: scripted determinism (LangGraph, Claude Code Workflow) is predictable but rigid, model-driven (Codex, Devin) is flexible but unpredictable, and hybrid (Windsurf 2.0) acts as a command center integrating multiple agents. The choice depends on how much predictability you need.

Pricing Watch | OpenAI Retires GPT-5.5 from ChatGPT/Codex on 10/14, API Untouched

OpenAI announced on 2026-09-14 that GPT-5.5 will retire from ChatGPT, ChatGPT Work, and Codex on 2026-10-14 — but calling `gpt-5.5` directly through the OpenAI API is unaffected. This is a product-surface retirement, not an API sunset. On the official pricing page, gpt-5.5's short-context input/output is $5.00/$30.00 per million tokens; the officially recommended Codex replacement, GPT-5.6 Sol, is $4.00/$20.00 — cheaper on both input (↓20%) and output (↓33%). The gap between announcement and shutdown is one month, far shorter than OpenAI's own documented minimum of six months' notice for GA models.

Model Card|GPT-6 Astra

GPT-6 Astra (API ID: gpt-6-astra): released by OpenAI on 2026-09-03, 1,050,000-token context window, 128,000 max output tokens, input $10.00 / output $50.00 per 1M tokens (cached input $1.00), closed-source; 99.9% on ARC-AGI-3 under OpenAI's own harness (62.7% on the standardized harness), 97.6% on FrontierMath Tier 4, 100% on ExploitBench; the first model rated 'Critical' cybersecurity capability under OpenAI's Preparedness Framework; yet scores only 61 on the neutral Artificial Analysis Intelligence Index — tied with predecessor GPT-5.6 Sol and behind Claude Fable 5.1's 66

Pricing Watch | OpenAI Assistants API Sunsets, Migration Forces a Model Choice

OpenAI's Assistants API (/v1/assistants, /v1/threads, /v1/threads/runs) officially sunset on 2026-08-26 — announced a year in advance, zero grace period, no automated migration tool. This isn't a pricing change on its own, but the forced migration also forces a model choice: workloads that ran on o3 ($2.00/$8.00 per million input/output tokens) via Assistants have no direct successor. OpenAI's official recommendation is GPT-5.6 Sol ($4.00/$20.00, cost ↑129%), but Terra ($2.00/$12.00, ↑29%) is often good enough in practice — a 44% gap between the two paths.

Looplane's ModelProvider multi-gateway: multiple protocols, one canonical contract

Looplane collapses OpenAI-compatible, Responses, Anthropic, Gemini, Workers AI, scripted, and experimental Codex OAuth adapters into one `ModelProvider` contract. The Codex OAuth transport reads SSE but still reduces it inside the adapter into one canonical `ModelTurn`; AgentRunner does not consume token deltas.

GPT——Closed API for Revenue, Open GPT-OSS for Ecosystem: the Unified Routing Platform Behind the World's Largest AI Service

GPT is OpenAI's LLM family, from 117M parameters in 2018 to the three-tier GPT-5.6 Sol/Terra/Luna lineup in 2026, serving 1B+ users and 2M enterprise customers. GPT-5.6 Sol leads LiveBench 81.1%, Terminal-Bench 2.1 88.8%, and Artificial Analysis Coding Agent Index 80 across multiple agentic benchmarks, while OpenAI's first open-weight model GPT-OSS ships under Apache 2.0.

Pricing Watch | OpenAI Cuts GPT-5.6 Sol Official Prices by 20-33%

OpenAI officially lowered GPT-5.6 Sol standard rates from $5.00/$30.00 to $4.00/$20.00 per million tokens (input/output; input ↓20%, output ↓33%), effective 2026-08-21, promotional period at least through 11/21. This is OpenAI's own price cut — not an OpenRouter/Cloudflare-style platform promo (see previous post). The two now stack: OpenRouter's 50% discount applies on top of the new $4/$20 base, yielding $2.00/$10.00.

Pricing Watch | GPT-5.6 Sol Half-Price on Both OpenRouter and Cloudflare Through 9/18

GPT-5.6 Sol standard rates through OpenRouter and Cloudflare AI Gateway drop from $5.00/$30.00 to $2.50/$15.00 per million tokens (input/output, -50%); Flex goes as low as $1.25/$7.50. Promo runs through 2026-09-18. Discount applies only to platform-managed billing (Unified Billing / non-BYOK) traffic — OpenAI's own API pricing is unchanged.

Agent Plugins 1.0: OpenAI, Google, and AWS Unite to Standardize AI Agent Extensions

Agent Plugins 1.0 is a packaging format that bundles Agent Skills (markdown instructions) and MCP server configs into a single directory, loadable by ChatGPT, Cursor, GitHub Copilot, Kiro, and VS Code. It's not a new protocol — it's the wrapper above protocols. Vercel initiated it, OpenAI/AWS/Microsoft/Cursor co-authored it, and Google joined on launch day. Anthropic isn't on the governance board, but MCP is a core primitive of the spec.

aideep-dive

The FDE War: Why OpenAI and Anthropic Are Both Copying Palantir's Playbook

MIT research says 95% of enterprise AI pilots yield zero return. OpenAI and Anthropic announced multi-billion-dollar joint ventures in the same week, wholesale adopting the Forward Deployed Engineer model that Palantir has used for over a decade to bring AI into the enterprise battlefield.

aideep-dive

OpenAI's Codex Secure Deployment Strategy: Sandboxing, Auto-review, and Enterprise Governance

In May 2026, OpenAI published its internal Codex deployment practices: sandboxes define technical boundaries, approval policies determine when to pause, Auto-review delegates approval decisions to a sub-agent instead of a human, and Managed configuration lets enterprise admins enforce policies top-down. The core philosophy: zero friction for low-risk actions, mandatory review for high-risk ones.

aideep-dive

OpenAI Workspace Agents: From Custom GPTs to a Team Automation Platform

On 2026/4/22 OpenAI launched Workspace Agents — powered by Codex, capable of long-running cloud execution, and integrating with Slack/Salesforce/Google Drive. They are the enterprise successor to Custom GPTs.

aiguide

Inside the Codex Agent Loop: How OpenAI Keeps AI Agents Iterating

A detailed look at OpenAI's Codex agent loop design: how prompts are constructed, how multi-turn conversations are managed, how prompt caching prevents cost explosions, and how context window auto-compaction works.

aiguide

Codex App Server: How OpenAI Turned an Agent Harness into a Universal Protocol

OpenAI wrapped the Codex harness as a JSON-RPC over stdio App Server, enabling VS Code, JetBrains, Web, and desktop apps to share a single agent loop. Three core primitives: Item, Turn, and Thread.

OpenAI Wrote 1 Million Lines of Code with Codex: Harness Engineering in Practice

An OpenAI internal team spent 5 months with 3 people and 0 lines of hand-written code, delivering a complete product using Codex. This article distills their core lessons on AGENTS.md design, repo-local knowledge bases, architecture enforcement, and entropy management.

aiguide

15 Agent Frameworks Worth Watching in 2026

Sorted by GitHub Stars, a survey of 15 mainstream AI Agent frameworks in 2026 — their positioning, key features, and ideal use cases. Not a ranking — it's a map.

Codex CLI: A Complete Guide to OpenAI's Open-Source Terminal Coding Agent

Codex CLI is OpenAI's open source terminal coding agent (Rust, Apache-2.0, ~106.6k stars) with MCP, subagents, image input, code review, and Skills. The model line is now GPT-5.6 Sol / Terra / Luna, and the desktop app, CLI, and IDE extension share one config.toml.

OpenClaw's Model Requirements and Provider Ecosystem: Provider, Model, and Runtime Are Three Different Things

OpenClaw's hard requirement for a model is tool use plus a large enough context — onboarding only auto-suggests a local model when it confirms tool support and at least a 16K context window. The easier thing to get wrong is that provider, model, and agent runtime are three separate layers: an `openai/*` ref does not mean Codex.