Two years ago, a senior engineer introduced Dapr to the company. It was a good decision. The team stopped writing broker-specific messaging code, services talked to each other securely, and new features shipped faster. That engineer has since moved on. Dapr is still running underneath dozens of services, but nobody on the current team has ever upgraded it, nobody is sure which version is in production, and a Slack message from last month asked, unanswered, “does anyone know when the Dapr certificates expire?”
This is one of the most common situations for teams that adopted Dapr early. If you’re an engineering manager, a tech lead or the “accidental platform engineer” at a small or mid-sized company, this article is a practical guide to taking ownership of a Dapr installation you inherited, without hiring a dedicated platform team.
Why inherited infrastructure is risky
Dapr is designed to fade into the background, which is a strength until nobody is paying attention. The risks of an unowned installation build up quietly:
- Version drift. Each Dapr release is supported for a limited period. Fall a few versions behind and you lose security fixes, and the eventual upgrade becomes a multi-step project.
- Certificate expiry. Dapr’s mTLS relies on root and issuer certificates that are valid for a long but finite period. If they expire, service-to-service calls fail across the cluster.
- Configuration drift. Components, resiliency policies and sidecar settings were set up for the traffic of two years ago, and nobody has revisited them.
- Knowledge loss. The reasons behind non-obvious settings left with the person who chose them.
None of these causes trouble on a normal Tuesday. All of them cause trouble on the worst possible day.
Step 1: Take inventory (one afternoon)
Start by finding out what you actually have. On each Kubernetes cluster:
dapr status -k
dapr mtls expiry -k
kubectl get components,configurations,resiliencies -A
Record:
- The control-plane version and whether it runs in high-availability mode
- The sidecar versions across workloads (they may differ if some pods haven’t restarted in a long time)
- The certificate expiry date
- Every component (state stores, pub/sub brokers, bindings, secret stores) and which apps use it
- Any resiliency policies and custom configurations
Put this in a single page in your team’s documentation. It’s probably the first time it has been written down.
Step 2: Remove the time bombs (first two weeks)
Two items deserve immediate attention because they can cause cluster-wide outages.
Certificates. If the expiry date is within six months, plan a renewal now. The Dapr CLI can generate and apply new certificates; practise in a non-production cluster first. Then put a calendar reminder and, better, an alert in place well ahead of the next expiry.
Unsupported versions. Check your control-plane version against Dapr’s supported releases. If you’re out of support, plan a stepwise upgrade through each minor version, upgrading the control plane first and then restarting workloads to pick up the new sidecars. Don’t try to jump several versions at once.
Step 3: Get visibility (first month)
You can’t own what you can’t see. The minimum useful set of signals for a small team:
- Control-plane pod health
- Sidecar injection failures
- Component initialization errors (usually a sign of a broken secret or connection string)
- Error rates and latency for service invocation and pub/sub per app
- Sidecar CPU and memory compared to their limits
- Days until certificate expiry
You can build this with Prometheus and Grafana, since Dapr exposes metrics from every sidecar and control-plane component. For a small team, though, building and maintaining dashboards is exactly the kind of work that gets dropped when priorities shift.
This is where purpose-built tooling earns its place. Diagrid, the company founded by Dapr’s original maintainers, offers a free SaaS tool called Dapr Ops Dashboard that connects to your clusters and provides much of this out of the box: prebuilt dashboards covering more than 150 metrics, automated control-plane upgrades and certificate rotation, and an advisor that checks your installation against more than 50 best practices for security, reliability and resource usage. One online retailer running more than 70 microservices on Dapr, processing over 80,000 orders a day, uses it so a small team can manage the platform without dedicated Dapr specialists.
Step 4: Write the runbooks you wish you’d inherited
Three short runbooks cover most of what a small team needs:
Upgrade runbook. Read the release notes, upgrade the non-production cluster, run integration tests, upgrade the production control plane, roll workloads in batches, verify versions. Schedule upgrades on a regular cadence, perhaps every second minor release, so they never become big projects.
Certificate runbook. How to check expiry, how to renew, how to verify that services can still communicate afterwards, and who gets alerted.
Incident runbook. What to check first when services can’t reach each other or messages stop flowing: control-plane health, sidecar logs, component errors, recent configuration changes.
Keep each one short enough to follow at 3 a.m. Test each one at least once.
Step 5: Spread the knowledge
The original problem was that Dapr knowledge lived in one person’s head. Don’t recreate it.
- Rotate upgrades between team members, so at least three people have done one.
- Add a short “how Dapr works here” section to engineering onboarding.
- Review Dapr configuration changes in pull requests, like any other infrastructure code.
- Keep the inventory page current, and update it as part of every upgrade.
When to get outside help
A small team can run Dapr well with the steps above. Outside help makes sense when:
- You’re several versions behind and the upgrade path looks risky
- Dapr runs business-critical workloads and you need guaranteed response times during incidents
- You’re planning a significant change, such as moving to Dapr Workflow, adding clusters or changing brokers, and want an expert review
- Your compliance requirements call for a supported, patched distribution
Commercial Dapr support from Diagrid is one option, typically combining enterprise support, upgrade guidance and access to the people who maintain the project. Whatever route you choose, decide before an incident whether you’ll need it.
A 30-day ownership checklist
- ☐ Inventory written: versions, HA mode, components, certificate expiry
- ☐ Certificate expiry more than six months away, with an alert configured
- ☐ Control plane on a supported version
- ☐ Control plane running in high-availability mode
- ☐ Basic dashboards and alerts for control-plane health, component errors and sidecar resources
- ☐ Upgrade, certificate and incident runbooks written and tested once
- ☐ At least two team members able to perform an upgrade
- ☐ A decision recorded on whether you need external support
Conclusion
Inheriting Dapr isn’t a problem. It usually means someone made a good architectural decision that has quietly paid off for years. The risk is in leaving it unowned. With an afternoon of inventory, a couple of weeks spent removing the time bombs, a month to build basic visibility, and a few well-tested runbooks, a small team can turn an orphaned installation into a well-run platform. And if you’d rather spend your engineers’ time on product work than on infrastructure, free Dapr operations tooling and expert support are there to carry part of the load.