Engineering Intelligent Operations Platforms
Autor Mateen Ali Anjumen Limba Engleză Paperback – 25 feb 2027
This book is a timely and authoritative guide for platform and DevOps engineers navigating the rapid convergence of AI and operations. As traditional responsibilities expand beyond infrastructure and CI/CD pipelines, engineers must now design and operate intelligent systems from anomaly detection pipelines and ML models to LLM-powered applications. This book provides the first unified framework that brings together AIOps, MLOps, and LLMOps, translating complex AI concepts into practical, production-ready strategies. Grounded in over a decade of real-world experience and reinforced by peer-reviewed research, it equips readers with the knowledge to build scalable, intelligent platforms using open, vendor-neutral tooling.
Structured across five comprehensive parts, the book progresses from foundational concepts to practical implementations. It begins with platform engineering fundamentals, OpenTelemetry-based observability, and AI-assisted infrastructure as code. It then dives into AIOps, covering anomaly detection, ML-driven FinOps, and AI-powered chaos engineering. The MLOps section walks through the complete model lifecycle pipelines, serving, monitoring, and drift detection, tailored specifically for platform engineers. The LLMOps section explores prompt management, RAG architectures, DevOps tooling powered by LLMs, and governance practices for secure AI systems. The final section integrates these disciplines into a unified platform architecture, complete with Kubernetes-based reference implementations, migration strategies, and organizational best practices. Each chapter includes hands-on examples in Python, Kubernetes, and Terraform, along with measurable benchmarks and reproducible projects.
By the end of this book, readers will be able to design, build, and operate a fully integrated intelligent operations platform that unifies AIOps, MLOps, and LLMOps. They will gain practical skills to deploy AI-driven systems at scale, implement observability pipelines as ML-ready data sources, manage model and LLM lifecycles in production, and evolve their organizations toward AI-first operations. This book empowers engineers to move beyond reactive DevOps toward autonomous, intelligent platforms that define the future of modern infrastructure.
What will you learn:
- Build ML-based anomaly detection pipelines from OpenTelemetry telemetry with benchmarked accuracy
- Deploy, serve, and monitor ML models on Kubernetes with automated lifecycle management
- Operate LLM applications using RAG, prompts, monitoring, and governance in production
- Apply AI techniques to FinOps, chaos engineering, IaC generation, and incident response
- Design unified platforms integrating AIOps, MLOps, and LLMOps on shared infrastructure layers
Who is it for:
This book targets DevOps engineers, platform engineers, and SREs with 3-10 years of experience who are already skilled in Kubernetes, Terraform, CI/CD, and basic Python, and are now taking on AI/ML workloads without formal training. It also supports engineering managers evaluating AIOps, MLOps, and LLMOps adoption. Readers are expected to have hands-on experience with Linux, containers, and at least one major cloud platform, while all required AI/ML concepts are taught in a practical, operations-focused context.
Preț: 394.45 lei
Preț vechi: 493.06 lei
-20% Precomandă
Carte nepublicată încă
Specificații
Ilustrații: Approx. 400 p.
Dimensiuni: 155 x 235 mm
Ediția:First Edition
Editura: APRESS L.P.
Colecția Apress
Notă biografică
Mateen Ali Anjum is a Staff DevOps Engineer with over 12 years of experience designing and operating production infrastructure at scale. He is the founder and principal consultant at Phono Technologies Inc., a DevOps consultancy based in Ontario, Canada, where he leads cloud architecture, Kubernetes platform engineering, and AI/ML operations projects for enterprise clients.
Mateen has authored six peer-reviewed research papers on topics that directly inform this book: self-healing infrastructure systems (Springer Journal of Cloud Computing), ML-driven FinOps (PeerJ Computer Science), platform engineering (Frontiers in Computer Science), AI-driven chaos engineering (Wiley Software: Practice and Experience), OpenTelemetry-based AIOps with empirical ML benchmarks (IEEE Access), and LLM-assisted Infrastructure as Code (Elsevier Information and Software Technology).
Cuprins
Part 1: Foundations.- Chapter 1: The Convergence of AI and Operations.- Chapter 2: Platform Engineering as the Foundation.- Chapter 3: Observability for Intelligent Systems.- Chapter 4: Infrastructure as Code in the AI Era.- Part 2: AIOps.- Chapter 5: Introduction to AIOps.- Chapter 6: Building Anomaly Detection Pipelines.- Chapter 7: FinOps Meets AIOps.- Chapter 8: AI-Driven Chaos Engineering.- Part 3: MLOps.- Chapter 9: MLOps Fundamentals for Platform Engineers.- Chapter 10: ML Pipeline Orchestration.- Chapter 11: Model Serving and Inference at Scale.- Chapter 12: ML Monitoring and Drift Detection.- Part 4: LLMOps.- Chapter 13: LLMOps: A New Operational Discipline.- Chapter 14: Building LLM-Powered DevOps Tools.- Chapter 15: RAG Architectures for Operations.- Chapter 16: Securing and Governing AI Systems.- Part 5: Putting It All Together.- Chapter 17: The Unified Ops Platform.- Chapter 18: The Future of Intelligent Operations.- Appendix A: Tool Comparison Matrix (AIOps/MLOps/LLMOps).- Appendix B: OpenTelemetry Quick Reference.- Appendix C: Terraform Modules for AI Infrastructure.- Appendix D: Further Reading and Research Papers.