Skip to content

Runbook

Internal operations guide for BANA - for developers, ops/SRE, and support staff. The deep mechanics live under Infrastructure; this runbook is the "what to do now" layer on top.

Internal use only

Operational procedures, incident playbooks, and support workflows. Do not share outside the NEXPANDO team.

Sections

SectionPurposePrimary audience
DeploymentDeploy, roll back, run migrationsDev, Ops
OperationsDay-to-day checks: services, queues, DB, cache, logsOps, On-call
IncidentsTriage + playbooks for outagesOn-call, Ops
SupportCustomer triage and common fixesSupport

Severity & on-call

SeverityMeaningResponse targetPath
S0Full outage - API down, DB lost, payments failing globallyacknowledge ≤ 15 min, all-handsOn-call → Incident Commander → Eng lead
S1Major degradation - one service down, high error rateack ≤ 30 minOn-call → service owner
S2Minor - one feature degraded, no money impactnext business dayservice owner
S3Cosmeticbacklog-

Escalation: On-call engineer → Incident Commander (rotating) → Engineering lead → CTO.

The on-call rotation, pager and IC roster are owned by the team and tracked outside the wiki (no contacts in the repo). Link the tool here once chosen.

Service levels (proposed - to ratify)

Proposed, not committed

BANA has no ratified internal SLOs yet. The targets below are a starting proposal for the team to agree - not commitments.

TargetProposedNotes
API availability99.5% / monthexcludes planned maintenance
Latency (p95)< 500 mscore read/write paths
RTO (recovery time)≤ 1 hourrestore service after S0
RPO (data loss)≤ 15 minDB backup cadence must meet this

Conventions

  • Severity - S0/S1/S2/S3 (above). Commands are copy-pasteable; placeholders use <angle-brackets>.
  • Keep entries short and recipe-style; link to Infrastructure for depth.

Proprietary and Confidential. Unauthorized copying, distribution, or use of this software is strictly prohibited.