15 prompts that cut a coding bill from $7,800 to $129¶
This article presents a cost optimization framework for AI coding workflows by splitting work between two models with distinct strengths. A team was spending $62,000/month running everything through Claude Opus 4.8. By routing orchestration to Opus and execution to Kimi K2.6 Agent Swarm, they cut costs to $7,800/month — an 88% reduction saving $54,200/month or $650,400/year, with the same or better output quality.
The core insight: strategy vs. execution¶
The central claim is that Claude Opus 4.8 and Kimi K2.6 are not interchangeable competitors but complementary layers. Opus 4.8 is positioned as a strategist: it plans, reasons, makes judgment calls, synthesizes conflicting information, and maintains quality standards. Anthropic's Dynamic Workflows feature was built so Opus can orchestrate hundreds of parallel subagents — not be those subagents. Kimi K2.6, scoring 58.6% on SWE-Bench Pro, is the executor: it runs 300 domain-specialized sub-agents in parallel, coordinates up to 4,000 steps, and produces real files (PDFs, spreadsheets, code). At $0.60 per million tokens versus Opus at $15 per million tokens, the economics are stark. The article's routing rule is simple: if you can write a rubric that a machine could grade, Kimi executes it. If you cannot write that rubric, Opus handles it.
15 prompts across three layers¶
The article organizes 15 prompts into three sections. The orchestration section (Prompts 1–5) covers Opus as project planner, dynamic task routing between models, quality rubric definition before execution, Opus reviewing Kimi's output against that rubric, and Opus assembling final deliverables from Kimi's components. The key discipline is that Opus never writes the actual content — it only produces the plan, the briefs, and the quality checks. Prompt 2 explicitly splits a task list into Group A (Opus-level: reasoning-heavy, ambiguous, high-stakes) and Group B (Kimi-level: execution-heavy, clear output spec, parallel-friendly). Prompt 4 creates a revision brief after review so Kimi can fix targeted issues without re-explanation.
The execution section (Prompts 6–10) translates Opus plans into Kimi project briefs, handles batch execution with exact output specs, runs research-to-deliverable in one pass, saves completed workflows as reusable Skills, and tracks costs per run. The Skill concept (Prompt 9) captures the full cycle — Opus orchestration prompt, Kimi brief template, review checklist, expected I/O — so recurring workflows start from a template rather than scratch. Prompt 10 logs actual token counts and costs against the counterfactual of all-Opus pricing.
The routing system section (Prompts 11–15) builds the ongoing decision infrastructure: weekly workflow audits to reclassify tasks between models, a routing decision tree for classifying incoming requests in 30 seconds, monthly cost optimization reviews, standardized handoff templates between the two models, and an ROI report for stakeholders. The cost comparison at the claimed scale: Opus handling 30% of workflow volume costs $18,600/month; Kimi handling the remaining 70% costs $1,240/month; versus $62,000/month for all-Opus. At smaller scale, the author's title references a personal bill dropping from $7,800 to $129.
The replacement economics¶
The article gives concrete before-and-after numbers for three workflow types. Fifty competitor landing page research briefs that would cost $25,000 in agency fees run for $4–6 in Kimi tokens. One hundred tailored outreach emails that would cost $2,000–5,000 from a copywriter execute in one sitting. A 30-codebase technical audit that would cost $15,000–40,000 in consultant time runs for $12–40 in tokens. The quality assurance mechanism is iterative: Opus defines the brief, Kimi executes at scale, Opus reviews against the pre-defined rubric, Kimi revises targeted failures, Opus approves. The article claims quality incidents remain the same or decrease because Opus reviews all final output, catching problems that a single-model workflow might miss precisely because the orchestrator is not also the executor. via model cli/openclaw/main
17s · cli/openclaw/main