Prompt: Create an operations playbook for a 50-person tech company in Johannesburg. Cover: daily/weekly/monthly rhythms, key meetings, escalation paths, and the 5 most important operational KPIs to track.
Operations Playbook – 50‑Person Tech Company
Location: Johannesburg, South Africa (UTC+02:00)
Audience: COO / Head of Ops, Engineering, Product, Finance & HR Leaders
> This playbook is a living document that can be stored in Confluence/Notion or any shared workspace. It assumes the company already uses common dev‑ops tooling – Jira/ClickUp for task tracking, GitHub/GitLab for code, PagerDuty/Datadog for alerts, and Slack/Teams for communication.
---
1. Daily Rhythm
| Time | Activity | Owner | Tools | Notes |
|------|----------|-------|-------|-------|
| 09:00 – 09:15 | Team Stand‑up (Daily Scrum) | Squad Leads (Engineering & Product) | Slack /standup or Zoom | • 2 min per person – What was done, what will be done, blockers. <br>• Record blockers in the sprint board and tag @ops. |
| 09:30 – 10:00 | Ops Pulse & Incident Review | Ops Manager | PagerDuty Dashboard + Datadog | • Check for new incidents (< 5 min response). <br>• Update incident tickets with status, ETA, stakeholder updates. |
| 12:30 – 13:00 | Finance/HR “Pulse” | Finance Lead & HRBP | Excel / Quickbooks / Workday | • Verify payroll runs, check for any UIF/PAYE exceptions. <br>• Capture any urgent head‑count changes (new hires or terminations). |
| 15:00 – 15:30 | Product Demo Prep | Product Owner | Notion / Figma | • Confirm demo content is ready; flag missing assets to design/engineering. |
| 16:30 – 17:00 | End‑of‑Day Wrap‑up (Ops) | Ops Manager | Teams/Slack | • Summarise day’s incident tickets, confirm SLA compliance. <br>• Post a short “daily status” to the channel (#ops-status). |
Daily SOP Highlights
| Step | Action | Responsible |
|------|--------|-------------|
| 1️⃣ Incident Triage | New alert → create PagerDuty ticket, add severity tag. | On‑call Engineer |
| 2️⃣ SLA Tracking | Update incident ticket with ETA; if > SLA, auto‑escalate to Ops Manager. | On‑call Engineer |
| 3️⃣ Ops Check‑list | Verify CI/CD pipelines passed, backup logs uploaded, security scan results OK. | Ops Technician |
| 4️⃣ Knowledge Capture | Log lessons learned in Wiki (link to incident). | Ops Lead |
---
2. Weekly Rhythm
> Weeks are numbered 1–52; the first Monday is “Week 1”.
| Meeting | Frequency | Time | Typical Attendees | Agenda | Owner |
|---------|-----------|------|-------------------|--------|-------|
| Sprint Planning | Every Mon‑Wed sprint cycle | 09:00 – 10:30 | Engineering, Product, QA | • Review backlog items.<br>• Set sprint goal.<br>• Estimate stories (Story Points). | Scrum Master |
| Engineering Sync | Tue & Thu | 11:00 – 12:00 | Eng. Leads + Ops Lead | • Architecture discussion<br>• Resource constraints<br>• Refactor / Tech debt backlog | Eng. Lead |
| Product Demo/Review | Wed | 14:00 – 15:30 | Product, Eng., Marketing | • Demonstrate new features.<br>• Collect stakeholder feedback.<br>• Approve for release. | Product Owner |
| Finance & HR Pulse | Thu (Friday morning) | 09:00 – 10:00 | Finance Manager, HRBP | • Head‑count report<br>• Burn rate update<br>• Payroll compliance checks | Finance Lead |
| All‑Hands | Friday 11:00 – 12:00 | All employees | CEO/COO, Ops, Eng, Product, Marketing, Sales | • KPI snapshot<br>• Company updates (product roadmap, hiring)<br>• Q&A | COO |
Weekly KPI Refresh
- Pull data from Jira & Datadog → refresh dashboard in PowerBI or Grafana.
- Send an automated email to the leadership team summarizing the 5 key KPIs (see Section 4).
---
3. Monthly Rhythm
| Meeting | Frequency | Time | Attendees | Purpose |
|---------|-----------|------|-----------|---------|
| Executive Review | First Monday of month | 09:00 – 11:00 | CEO, COO, CFO, VP Engineering, VP Product | • Deep dive KPI trends.<br>• Capital allocation & runway planning.<br>• Strategic decisions (new market, product line). |
| Ops Health Check | Mid‑month (15th) | 14:00 – 15:30 | Ops Lead + Eng. Leads + Finance Manager | • Incident trend review<br>• Vendor performance<br>• Budget vs spend analysis. |
| Vendor & Partner Review | Last Thursday of month | 10:00 – 11:30 | Procurement, Ops Lead, VP Engineering | • KPI on delivery lead‑time, quality score.<br>• Negotiate contract adjustments if needed. |
Monthly SOPs
Trigger: ≥ 3 incidents in a month or any critical incident.
Owner: Ops Manager + Eng. Lead.
Deliverable: 10‑minute root‑cause report + action items.
Trigger: End of each calendar month.
Owner: Finance Lead.
Process: Project next 3 months burn rate, runway; adjust headcount plan accordingly.
---
4. Escalation Paths
> Use a 3‑tier model: Level 1 – Ops / Eng. Team → Level 2 – Operations Manager → Level 3 – COO/Executive.
| Issue Type | Threshold | Who Escalates? | How |
|------------|-----------|----------------|-----|
| Service Outage > 5 min | Any system with SLO < 99.9% fails | On‑call Engineer | PagerDuty auto‑escalate to Ops Manager (3 min). |
| Security Breach / Data Leak | Any detected or suspected | Security Lead | Notify COO immediately via Slack + email; initiate incident response playbook. |
| Financial Exception > ZAR 100k | Payroll/Invoice variance | Finance Analyst | Flag in ticketing system, call Ops Manager for mitigation. |
| Product Release Delay > Sprint Target by 2 days | Feature not delivered | Product Owner | Escalate to VP Engineering & COO if impact on revenue. |
| Vendor SLA Violation (≥ 10%) | Delivery/Quality issue | Procurement Lead | Raise ticket to Vendor Manager, involve Ops Manager. |
> Escalation Flowchart (text)
``
Level 1 → Level 2 → Level 3
↓ ↓ ↓
On‑call → Ops Lead → COO/VPs
`
- Every escalated issue must create a ticket in Jira Service Desk with @ops
, @security or other relevant tags.
- The ticket should capture: Impact, Root cause hypothesis, Time to resolution target.
---
5. Five Most Important Operational KPIs
| KPI | Definition | Data Source | Frequency | Threshold / Target |
|-----|------------|-------------|-----------|--------------------|
| 1️⃣ Incident Mean Time to Recovery (MTTR) | Avg time from incident alert to full service restoration. | PagerDuty + Datadog. | Weekly & Monthly | < 30 min for critical services; < 15 min for core APIs. |
| 2️⃣ Release Cycle Time | Days between commit and production release. | GitHub Actions + Jira. | Weekly & Monthly | ≤ 3 days from sprint start to deploy. |
| 3️⃣ System Uptime / SLO Compliance | % of time services meet defined SLOs (e.g., 99.9%). | Datadog, NewRelic. | Daily Snapshot + Monthly Review | ≥ 99.5% overall; critical services ≥ 99.9%. |
| 4️⃣ Burn Rate & Runway | Cash spent per month vs cash on hand. | Finance ERP / Quickbooks. | Monthly | < R1 m per month burn → 6‑month runway. |
| 5️⃣ Employee Net Promoter Score (eNPS) | Employee engagement metric; gauge internal health. | Quarterly Survey (SurveyMonkey). | Quarterly | ≥ 30 points higher than industry baseline (~15–20). |
KPI Dashboard Example
- Tool: Grafana + PowerBI
- Visuals: Line chart for MTTR trend, bar chart for Release Cycle Time, gauge for Uptime, line for Burn Rate, single value eNPS.
- Alerting: Trigger Slack alert if MTTR > 45 min or uptime falls below target.
---
6. Sample Meeting Templates
Daily Stand‑up (2 Min per person)
`
[Name] – What did I finish yesterday?
[Name] – What am I working on today?
[Name] – Any blockers? (Escalate to @ops if critical)
``
Weekly Sprint Planning Agenda
| Time | Item |
|------|------|
| 0‑5 min | Review sprint goal & capacity. |
| 5‑20 min | Backlog refinement (story re‑prioritisation). |
| 20‑35 min | Estimate new stories with Pomodoro. |
| 35‑45 min | Identify dependencies & assign owners. |
| 45‑50 min | Confirm definition of done & acceptance criteria. |
| 50‑55 min | Q&A; blockers list for Ops/DevOps. |
Monthly Executive Review (PowerPoint skeleton)
- Opening – Vision recap, major wins.
- Financial Health – Burn rate, runway, cash flow chart.
- Operational Metrics – MTTR trend, uptime, release cycle time.
- Product & Market – Release roadmap, NPS/CSAT trends.
- People – Headcount, churn, eNPS score.
- Risks & Mitigations – Vendor SLAs, compliance items.
- Action Items & Decisions – Owner + due date.
---
7. Tool Recommendations (South African Context)
| Category | Tool | Why |
|----------|------|-----|
| CI/CD | GitHub Actions (free tier) | Global build agents, no infrastructure. |
| Issue Tracking | Jira Cloud | Scalable for sprint planning, integration with Confluence. |
| Incident Management | PagerDuty + Datadog | SLA management, auto‑escalation. |
| Ops Automation | Cloudflare Workers + D1 | Low‑cost edge compute; easy to host infra scripts (e.g., daily backup). |
| Finance | QuickBooks Online SA or Xero | Payroll integration with PAYE & UIF automatically calculated. |
| Survey | SurveyMonkey / Google Forms | eNPS collection in ZAR currency. |
> Local compliance:
> • Make sure payroll software calculates PAYE, UIF, and Covid‑19 tax credits correctly.
> • Maintain SARB and VAT reports monthly; store in a dedicated SharePoint folder.
---
8. Continuous Improvement Loop
- Review KPI trends – weekly to spot anomalies.
- Monthly Retrospective – each team reviews successes & blockers.
- Update SOPs – every 6 months or after major incidents.
- Training Sessions – quarterly on new tools (e.g., GitHub Actions best practices).
By anchoring your operations around the rhythms above, you give every function a predictable cadence, clear escalation paths, and a data‑driven focus that scales as the company grows beyond 50 people.
---
Prepared by Tom, Fractional COO at 2nth.ai – operational excellence for tech startups.