Cantitate/Preț
Produs

AI Cost Playbook

Autor Caio Incau
en Limba Engleză Paperback – 29 sep 2026
Bring FinOps discipline to production LLM systems by measuring AI spend, optimizing token usage, and building cost controls that scale with your applications
Key Features:
- Build tokenwatch to measure, attribute, and optimize production LLM costs
- Reduce AI spend using caching, batching, model routing, and RAG optimization
- Establish budgets, quotas, forecasting, and accountability for sustainable AI
Book Description:
LLM applications introduce a new kind of cloud economics. Costs are usage-based, model-dependent, and influenced by everything from prompt length and caching to RAG pipelines and autonomous agent loops. As grows, engineering teams need more than isolated cost-cutting tricks; they need FinOps for LLM systems.
The AI Cost Playbook shows you how to apply financial accountability and engineering discipline to production AI. You'll build tokenwatch, a Python-based cost observability and optimization toolkit, while learning to meter tokens, attribute spend, and calculate costs per request, user, and feature.
You'll then reduce unnecessary spend through prompt optimization, caching, batch processing, model routing and cascades, RAG optimization, and agent cost controls. You'll also evaluate the break-even economics of self-hosting and learn how to forecast future AI expenditure.
Finally, you'll turn optimization into an operating discipline by establishing budgets, quotas, forecasting, and cost accountability. With configurable pricing rather than hardcoded model costs, the techniques remain useful as providers, models, and pricing evolve.
What You Will Learn:
- Measure token usage and attribute LLM costs accurately
- Calculate AI costs per request, user, and product feature
- Optimize prompts to reduce unnecessary token consumption
- Use prompt caching and calculate its break-even point
- Cut workload costs with batching and model routing
- Optimize RAG pipelines and control agent-related costs
- Evaluate the economics of APIs versus self-hosted models
- Build budgets, quotas, forecasts, and AI cost governance
Who this book is for:
This book is for AI engineers, ML engineers, software engineers, platform engineers, technical leads, engineering managers, architects, and FinOps professionals responsible for building or operating LLM applications. It will also benefit technology leaders responsible for AI infrastructure and API spending who want to understand the economics behind production AI systems and establish better cost controls. Familiarity with LLM applications and basic Python will help readers get the most from the implementation-focused sections.
Table of Contents
- The AI Bill Shock
- Token Economics 101
- Measuring First: Token Metering and Cost Attribution
- Unit Economics: Cost per Request, User, and Feature
- Prompt Engineering for Cost
- Prompt Caching: Mechanics, Hit Rates, and Break-Even
- Batch Processing and Async Workloads
- Model Routing and Cascades
- RAG Cost Optimization
- Agents: The Cost Multiplier
- Self-Hosting Break-Even
- Governance: Budgets, Quotas, and Cost Accountability
- Negotiating and Forecasting
- Final Project: Full Tokenwatch Integration
Citește tot Restrânge

Preț: 16851 lei

Preț vechi: 21063 lei
-20% Nou

Puncte Express: 253

Carte tipărită la comandă

Livrare economică 14-28 noiembrie

Livrare prin curier în România Termenul estimat este afișat lângă disponibilitate.
Transport gratuit de la 40000 lei Plată online sau ramburs, în funcție de opțiunile comenzii.
Retur gratuit în 14 zile Comandă securizată și suport în română.

Specificații

ISBN-13: 9781808822155
ISBN-10: 1808822153
Pagini: 206
Dimensiuni: 191 x 235 x 11 mm
Greutate: 0.4 kg
Editura: Packt Publishing

Notă biografică

Caio Incau is an Engineering Manager with experience leading software engineering teams at scale. In his day-to-day work, he combines people leadership with deep technical expertise to deliver products that impact millions of users. He started using Claude Code out of curiosity, became an advocate after seeing the productivity gains firsthand, and wrote this book so other developers wouldn't have to figure everything out on their own.