Engineering

Incident Response Board

Track incidents from first alert to postmortem in one live board

The Incident Response Board gives engineering teams a single place to manage the full lifecycle of a production incident: intake, diagnosis, mitigation, monitoring, and resolution with follow-ups. Instead of scattering updates across chat threads and docs, every incident becomes one task that moves through five clear columns as work progresses. Paired with the MCP server, an AI on-call agent can file new incidents the moment an alert fires, move tasks as investigation and mitigation happen, and log root-cause notes directly on the task, so the board stays accurate in real time without manual busywork. Use it standalone via the kanban board with keyboard shortcuts and command palette, or drive it programmatically through the REST API or MCP tools.

kanboard.io/templates/incident-response

Template preview

Incident Response Board

5 columns
Triage

Newly reported incidents land here immediately so on-call can assess severity and assign an owner.

2

Payments API returning 500 errors

Spike in 5xx responses on /v1/charge starting 14:02 UTC. Sev-1, on-call paged, awaiting first responder ack.

Elevated login failures across web app

Auth service error rate at 12%, user reports coming in via support. Needs severity assessment and owner.

Diagnosing

Incident is actively being investigated to find root cause; owner is confirmed and gathering data.

2

Checkout latency spike investigation

p95 latency up 4x since deploy at 13:40 UTC. Pulling traces and comparing to last known-good deploy hash.

Intermittent 502s from load balancer

Reviewing LB access logs and upstream health checks to isolate whether it's a single backend node.

Mitigating

Root cause understood or workaround identified; a fix, rollback, or failover is actively being applied.

2

Rolling back checkout service to v2.14.1

Root cause confirmed as bad deploy. Rollback in progress via deploy pipeline, ETA 10 min.

Failing over primary DB to read replica

Applying failover runbook step 4-7 to relieve write contention on primary.

Monitoring

Mitigation applied; team is watching metrics to confirm stability before declaring the incident resolved.

2

Watching error rate post-rollback

Rollback deployed 14:20 UTC. Error rate dropping, holding 15-min monitoring window before closing.

Confirming replica failover stability

Write latency back to baseline. Watching for replication lag before marking resolved.

Resolved & Follow-up

Incident confirmed closed; captures postmortem action items and preventive follow-up work.

2

Postmortem: checkout deploy rollback

Incident closed 14:35 UTC. Follow-up: add pre-deploy latency canary check, owner assigned, due next sprint.

Follow-up: add replica alert threshold

Add alert for replication lag > 5s to catch similar DB incidents earlier.

How to run this board

Step 1

An alert or report triggers a new task filed in Triage with severity, affected system, and initial owner noted in the description.

Step 2

On-call moves the task to Diagnosing once they begin investigating, updating the description with findings as they emerge.

Step 3

When a fix or workaround is being applied, the task moves to Mitigating; once deployed, it shifts to Monitoring to confirm stability.

Step 4

After a stable monitoring window, the task moves to Resolved & Follow-up, where postmortem notes and action items are logged.

Step 5

Stale Resolved & Follow-up tasks are periodically cleaned up via delete_task to stay within plan task limits.

AI agent usage

Run it with AI agents

Open MCP setup

This playbook lets an AI agent connected via the Kanboard MCP server manage a full incident lifecycle: intake new alerts, run diagnostics, track mitigations, and close the loop with follow-up actions. The agent should keep exactly one task per incident, move it across columns as status changes, and log timestamps/owners in the task description as it works so the board stays the single source of truth during an active incident.

## Incident Response Board
Columns: Triage, Diagnosing, Mitigating, Monitoring, Resolved & Follow-up.
- Use `create_task` to file a new incident in Triage as soon as an alert fires; include severity, affected service, and first responder in the description.
- Use `move_task` to advance a task to Diagnosing once someone starts investigating, and to Mitigating once a fix or workaround is being applied.
- Use `update_task` to append findings, root cause notes, and mitigation steps directly to the task description or comments as they happen.
- Move tasks to Monitoring after a mitigation is deployed but before it's confirmed stable; move to Resolved & Follow-up only after the incident is confirmed closed, and list any post-incident action items in that task.
- Respect the plan limit (25 tasks on free) by archiving or deleting stale Resolved & Follow-up tasks with `delete_task` once follow-ups are done.
File a new P1 incident for the payments API returning 500s and put it in Triage with on-call owner assigned.
Move the database latency incident from Diagnosing to Mitigating and add a note that we're failing over to the replica.
List all tasks currently in Monitoring so I can check which incidents still need a stability window before closing.

Learn more about the board-centric workflow in Kanban for AI Agents or open the MCP guide.

Frequently asked questions

How many incidents can I track on the free plan?

The free plan supports 1 project and up to 25 tasks total, so you can run this board with roughly 25 active or recently resolved incidents at once. Delete or archive old Resolved & Follow-up tasks to stay under the limit.

Can I run multiple incident boards for different teams?

The free plan is limited to 1 project. If you need separate boards per team or service, upgrade to Pro ($12/mo or $96/yr), which supports multiple projects.

How does an AI agent update the board during an incident?

Using the MCP server, an agent can call create_task to file new incidents, move_task to shift them between Triage, Diagnosing, Mitigating, Monitoring, and Resolved & Follow-up, and update_task to log findings and mitigation notes as the incident progresses.

Does this board support automated alerting or paging?

No. The board itself only provides kanban board, columns, tasks, and MCP/REST API access for updating task state; it does not include built-in paging or alerting features. You'd trigger task creation from your existing alerting tool via the API or MCP server.

Can I access this board outside the browser?

Yes, you can manage tasks via the REST API, the MCP server for AI agents, or the macOS app, in addition to the standard kanban board interface with keyboard shortcuts and command palette support.

What happens to follow-up tasks after an incident is resolved?

Follow-up action items live as tasks in the Resolved & Follow-up column. Once those action items are completed, delete them with delete_task to free up space against your plan's task limit.

Related templates

Ready to create this board?

Sign in and use the template to create a project with columns and sample tasks.

Sign in to activate