Skip to content
All tags

#token-budget

4 posts
tech

Codex Turn 狀態機:TurnContext、StepActivation、Context Manager 與壓縮觸發

TurnContext 在 turn 初始化時捕獲所有設定(模型、審批、token budget),後續 step 透過 StepContext 讀取快照;StepActivation 驗證設定變更不違反 legacy 安全約束;ContextManager 用 Arc<Vec> + 版本號實現 Copy-on-Write 歷史共享;壓縮觸發條件為 token_remaining < threshold,支援 remote v1/v2、local、model fallback 四條路徑。

aideep-dive

15 Walls for Building Your Own Auto-Dev Agent: Concrete Lessons from Stripe Minions

Stripe Minions says 'The walls matter more than the model,' but the case studies from four Silicon Valley companies never explained how to actually build those walls. This post breaks down the 15 walls we implemented in the daodao auto-dev agent: what each wall prevents, where the files live, and what the tradeoffs are. Tier 1 is mandatory, Tier 2 strengthens governance, Tier 3 is serious governance.

RAG Cost Optimization: Minimizing the Cost of Every Query

RAG system costs come from LLM tokens, Embedding APIs, and vector search. Every stage has room for cost reduction, but you need to verify that optimizations don't sacrifice too much quality.

RAG Quota System: Controlling LLM Costs with Dual Limits

Limiting request count alone is not enough — a single long query can consume ten times the tokens of a normal one. Dual quotas (request count + token count) are what truly control costs.