sources · source-BATCH-2026-003-002
Building Secure and Reliable Systems: Best Practices for Designing, Implementing, and Maintaining Systems
A multi-author practitioner volume integrating security and reliability across design, implementation, deployment, incident response, recovery, organization, and culture.
Open the canonical original source
Source at a glance
- Source type
- book
- Publisher
- O’Reilly Media, Inc.
- Source or access date
- 2020-04-08
- Analysis depth
- complete edition deep analysis
- Rights treatment
- link and paraphrase
- Human review
- Murray Newlands · 2026-08-16
Important limitations
- Many recommendations reflect Google's scale, staffing, engineering maturity, and ability to build custom control planes.
- The chapter evidence is primarily practitioner experience rather than externally replicated causal analysis.
- The book predates modern general-purpose AI agents and does not define sponsor, delegation, session, model, tool-purpose, or decision-provenance identity fields.
- Privacy, legal process, and geography are acknowledged but not developed into jurisdiction-specific operating models.
Source-located statements
18 reviewed statements are indexed from this source.
- Security and reliability should be designed and operated together across the system lifecycle because late-stage controls cannot reliably compensate for unsafe architecture and operations. (Preface > Why We Wrote This Book)
- Security and reliability are emergent system properties with common needs for prevention, detection, containment, and recovery, even though adversarial intent changes the threat model. (Chapter 1 > Reliability and Security Commonalities)
- Risk assessment should model capable adversaries, valued assets, likely attack paths, and the consequences of successful attacks. (Chapter 2 > Risk Assessment Considerations)
- Design decisions should make security, reliability, usability, cost, and other nonfunctional requirements explicit so tradeoffs are evaluated rather than hidden. (Chapter 4 > Design Objectives and Requirements)
- Least privilege applies to people, automated jobs, and machines: each principal should receive only the access needed for the current responsibility. (Chapter 5 > Least Privilege)
- Organizations should classify access by the potential impact of misuse and apply controls proportionate to the resulting risk tier. (Chapter 5 > Classifying Access Based on Risk)
- Sensitive operations should be exposed through narrow functional APIs and recorded in action-oriented audit logs rather than granted as broad administrative access. (Chapter 5 > Small Functional APIs; Auditing)
- Authorization policy can evaluate the action, arguments, request source, principal metadata, and server-side context rather than treating a credential as sufficient permission. (Chapter 5 > A Policy Framework for Authentication and Authorization)
- High-impact access should use controls such as multi-party approval, temporary grants, structured business justification, and monitored breakglass procedures. (Chapter 5 > Advanced Controls)
- Identities, authentication flows, interfaces, trusted computing bases, and security boundaries should be understandable enough for operators to reason about normal and failure behavior. (Chapter 6 > Understandable Identities, Authentication, and Authorization)
- Automation should be layered with independent checks, bounded failure domains, progressive rollout, and retained human intervention so a flawed action cannot spread without limit. (Chapter 8 > Automate Responsibly; Controlling the Blast Radius)
- Recovery requires explicit revocation, knowledge of intended state, tested restoration paths, and mechanisms that do not depend on the component already known to be compromised. (Chapter 9 > Use an Explicit Revocation Mechanism; Know Your Intended State, Down to the Bytes)
- Deployment systems should verify artifacts and their provenance at controlled choke points instead of relying only on the identity of developers or on earlier review events. (Chapter 14 > Verify Artifacts, Not Just People; Binary Provenance)
- Security logs should resist alteration while their collection, access, retention, and content are governed to avoid creating unnecessary privacy exposure. (Chapter 15 > Design Your Logging to Be Immutable; Take Privacy into Consideration)
- Disaster response should be rehearsed through structured exercises and, where safe, production tests because an untested recovery plan is only an assumption. (Chapter 16 > Disaster Planning)
- Major incidents benefit from an explicit incident commander and separate coordination, mitigation, investigation, and communication responsibilities. (Chapter 17 > Crisis Management)
- Security and reliability are responsibilities of everyone who changes or operates the system, supported by specialists, embedded champions, clear ownership, and executive and board stakeholders. (Chapter 20 > Who Is Responsible for Security and Reliability?)
- Leaders should deliberately shape incentives, treat failure as inevitable, reward learning and reporting, and provide sustainable resources for security and reliability work. (Chapter 21 > Culture of Inevitability; Culture of Sustainability; Align Project Goals and Participant Incentives)
Evidence lineage and transparency
| Relationship field | Linked identifiers |
|---|---|
| author ids | person-BATCH-2026-002-019, person-BATCH-2026-002-020, person-BATCH-2026-002-021, person-BATCH-2026-002-022, person-BATCH-2026-002-023, person-BATCH-2026-002-024 |
Machine review: ready for human review. Human review: approved. Workflow: published.
Complete structured record
- source id
- source-BATCH-2026-003-002
- canonical title
- Building Secure and Reliable Systems: Best Practices for Designing, Implementing, and Maintaining Systems
- alternate titles
- Building Secure & Reliable Systems
- source type
- book
- series or parent source
- Unknown
- publisher
- O’Reilly Media, Inc.
- channel
- Unknown
- speaker ids
- author ids
- person-BATCH-2026-002-019, person-BATCH-2026-002-020, person-BATCH-2026-002-021, person-BATCH-2026-002-022, person-BATCH-2026-002-023, person-BATCH-2026-002-024
- institutional author
- Unknown
- organization references
- recorded at
- Unknown
- event date
- Unknown
- published at
- 2020-04-08
- updated at
- Unknown
- duration seconds
- Unknown
- language
- lang-en
- translated title
- Unknown
- translation method
- Unknown
- geography of speaker
- geography of organization
- United States
- geography discussed
- Global large-scale technology operations
- study geography
- original url
- https://google.github.io/building-secure-and-reliable-systems/
- canonical url
- https://www.oreilly.com/library/view/building-secure-and/9781492083115/
- archived url
- Unknown
- embed url
- Unknown
- doi
- Unknown
- canonical identity status
- Exact first online edition verified by title, ISBN, publisher, publication date, pagination, official content manifest, and normalized text hash
- repost status
- Unknown
- original source id
- Unknown
- rights status
- link_and_paraphrase
- ownership status
- third_party
- relationship to off
- Unknown
- transcript status
- Unknown
- transcript source
- Unknown
- transcript republication permission
- Unknown
- chapter markers
- [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object]
- analysis basis
- Complete official first-edition HTML: 35 files and 176,914 normalized words read in full
- topics
- topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity, topic-enterprise-infrastructure
- executive roles
- CISO, CIO, CTO, CEO, General Counsel, Board Director
- original abstract
- A multi-author practitioner volume integrating security and reliability across design, implementation, deployment, incident response, recovery, organization, and culture.
- inclusion rationale
- Provides governance patterns for least privilege, contextual authorization, artifact provenance, auditability, recovery, and organizational accountability that can be carefully adapted to AI agents.
- source quality dimensions
- {"authority":"Primary practitioner source from named Google and industry contributors","completeness":"Complete exact edition","independence":"Moderate to low; many examples concern the authors' employer","reproducibility":"High; stable official chapter and section URLs"}
- methodology quality
- {"design":"Multi-author practitioner synthesis","causal_strength":"Low to moderate for design reasoning; low for generalized outcomes","transparency":"Tradeoffs are commonly stated; case selection is not systematic"}
- study design
- Practitioner guidance, technical examples, incident narratives, and case studies
- sample
- {"chapters":22,"content_files":35}
- population
- Large-scale software and infrastructure organizations
- date range
- {"edition":"2020"}
- funding
- Unknown
- sponsor
- peer review status
- Editorially reviewed technical book; no formal academic peer review identified
- findings
- limitations
- Many recommendations reflect Google's scale, staffing, engineering maturity, and ability to build custom control planes., The chapter evidence is primarily practitioner experience rather than externally replicated causal analysis., The book predates modern general-purpose AI agents and does not define sponsor, delegation, session, model, tool-purpose, or decision-provenance identity fields., Privacy, legal process, and geography are acknowledged but not developed into jurisdiction-specific operating models.
- correction ids
- retraction status
- Unknown
- content hash
- 6bf5050c08be0bced81767ac14bd9b3303e34b3efb667806b3c9e9d6b0ed8cf1
- accessed at
- 2026-08-15
- verification status
- machine_verified_human_approved
- source depth
- complete_edition_deep_analysis
- publication status
- published
- workflow status
- published
- machine review status
- ready_for_human_review
- human review status
- approved
- reviewed by
- Murray Newlands
- reviewed at
- 2026-08-16T23:17:22Z
Provenance and revision history
{
"provenance": [
{
"source_url": "https://google.github.io/building-secure-and-reliable-systems/",
"accessed_at": "2026-08-15",
"retrieval_method": "Complete exact edition retrieved from an official public source and reviewed in full",
"exact_locator": "All 35 official HTML files; chapter/section anchors",
"content_hash": "6bf5050c08be0bced81767ac14bd9b3303e34b3efb667806b3c9e9d6b0ed8cf1",
"batch_id": "BATCH-2026-003",
"prompt_id": "OEII-BOOK-DEEP-ANALYSIS",
"prompt_version": "2.0",
"notes": "Manifest hash ae702b4e64364217f179317b4a46de6a5e7c51045a79efa67522245d76667e16; no direct quotations stored"
}
],
"revision_history": [
{
"changed_at": "2026-08-16T23:17:22Z",
"changed_by": "Murray Newlands",
"summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
"batch_id": "BATCH-2026-003"
}
]
}