Kołczan Książek
Załóż konto
SRE for AI Systems1 / 1
Wydanie · Softcover

SRE for AI Systems

PNPravin Nair
—
rok wydania
308
stron
polski (Polska)
język
Softcover
typ okładki
Wydawca
BPB Publications (Z chęcią przeczytam książkę w języku polskim)
ISBN-13
978-93-7854-419-4
EAN
9789378544194
Data wydania
—
Książka
SRE for AI Systems · 1 wydanie
Wydawca o tym wydaniu

Description As AI and ML rapidly power modern digital services from recommendation engines to generative models, moving these workloads into production exposes critical gaps in traditional operations. Site reliability engineering (SRE) applies software engineering principles to infrastructure, making SRE uniquely positioned to own the reliability, observability, and resilience of complex AI-driven environments. This book bridges the gap between traditional SRE practices and the innovative strategies required to manage AI-driven infrastructures such as LLM models effectively, offering readers a focused guide for this transformative era. Each chapter walks the reader through core concepts such as versioning of models and data, pipeline monitoring, and security testing for AI APIs, while providing concrete SRE patterns for uptime, rollback, and incident management in AI-driven environments. You will learn how to design service-level objectives that reflect AI-specific quality metrics, implement robust monitoring and alerting for model drift and data quality, secure AI-backed APIs against prompt injection and other attacks, and coordinate releases across code, data, and models in a complex environment. By the end of this book, as an SRE engineer, you will be equipped to design, operate, and scale AI systems with the same rigor that you applied to traditional infrastructure. You will gain skills in building observable AI pipelines, managing versioned data and models at scale, securing AI APIs, and applying SRE principles such as error budgets, incident response, and automation to complex ML and generative AI workloads. What you will learn Applying SRE principles to AI and ML workloads. Monitoring AI pipelines for model quality, data drift, and output correctness. Defining SLIs and SLOs that reflect business metrics. Building resilient deployment and rollback strategies for LLM models. Equipping SRE teams to own AI pipelines and ML lifecycle. Building and scaling SRE teams to manage AI systems. Who this book is for This book is for site reliability engineers, ML engineers, software developers, and SRE managers supporting AI workloads. Readers should possess basic cloud infrastructure knowledge, a foundational understanding of software development, and familiarity with general machine learning principles. Table of Contents 1. Introduction to SRE and AI Systems 2. Reliability Challenges in AI Workloads 3. Reliability Failures in AI Systems 4. Monitoring and Observability for AI Systems 5. AI-enhanced Automation in SRE 6. Resilient Architecture in Cloud 7. Incident Management and Root Cause Analysis with AI 8. SLOs and Error Budgets for AI Systems 9. CI/CD and Testing Using AI and SRE Principles 10. Building and Scaling AI SRE Teams 11. Cultural Shifts and Collaboration in AI SRE 12. Future Trends and Ethical AI in SRE