Back to projects

Internal Analytics Dashboard for Ops & Support Teams

Designed and led the technical plan for an in-product analytics dashboard, pulling our ops and support teams out of Grafana, Zendesk, and ad-hoc SQL and into one place.

NestJSVue.jsPostgreSQLRedisZendesk API

01.Problem Statement

Our internal teams had to jump between three different tools just to answer basic questions: Grafana for operational metrics, Zendesk for support tickets, and someone running manual SQL queries whenever a question didn't fit either tool. There was no single place to see stalled orders, shipment exceptions, and support ticket trends side by side, which made daily ops reviews slower than they needed to be.

02.Architecture Overview

I led the technical design for a new Alert, Ops, and CS tab inside our existing admin dashboard, backed by our backend service. Most of the new metrics are served from a lightweight cache that refreshes every few minutes, so the dashboard stays fast without hammering the database. Support ticket data gets synced from Zendesk into our own database on an hourly job, since pulling it live on every page load would have hit Zendesk's rate limits fast.


   [ Admin Dashboard ] --> [ Alert / Ops / CS Tabs ]
           |
           v
   [ Backend API ] --> [ Cache Layer ]
           |                  |
           v                  v
   [ Orders / Shipments DB ]  [ Synced Support Tickets ]
                                      ^
                                      |
                              [ Hourly Zendesk Sync ]
    

03.Database Design

Added a dedicated table that mirrors the support ticket data we care about, kept in sync with Zendesk on a schedule and indexed for fast lookups by date and by client, so the dashboard can answer common questions without ever calling Zendesk directly.

04.Key Decisions & Tradeoffs

Decisions

  • Chose to sync support ticket data into our own database on a schedule instead of calling Zendesk live, so the dashboard stays fast and never gets rate-limited.
  • Reused the stalled-order and anomaly detection logic we already had in our alerting system, instead of writing a second version of the same business rules just for the dashboard.
  • Rolled every new tab out behind a single feature flag so we could get it in front of real ops and support users early and adjust before a full launch.
  • Made sure every number behind the dashboard respects the same client-level data boundaries our admin roles already enforce, so nobody sees data they shouldn't.

Tradeoffs

  • Accepted that support ticket numbers can be up to an hour old in exchange for a dashboard that always loads fast and never gets throttled by Zendesk.
  • Kept the first version focused on the metrics teams actually asked for, intentionally leaving fancier stuff like SLA percentiles for a later phase.

05.Scaling Considerations

Because the heavy lifting (ticket syncing, metric aggregation) happens in the background rather than on every page view, the dashboard stays responsive even as ticket and order volume grows. The caching strategy means most page loads never touch the database directly.

06.Failure Scenarios & Mitigation

  • Zendesk is down or slow: the sync job just waits and retries later; the dashboard keeps showing the last good data instead of breaking.
  • A metrics query fails temporarily: the dashboard falls back to the last cached result rather than showing an error to the ops team.

07.Engineering Challenges

  • Balancing freshness against cost: some numbers needed to feel close to real-time, while others were fine as a snapshot refreshed every few minutes. I designed different caching rules for each so teams get speed where it matters and accuracy where it counts.
  • Keeping support ticket data useful even when Zendesk itself was slow or temporarily unavailable, by always falling back to the last successful sync instead of showing errors.

08.Implementation Details

09.Impact & Outcome

Impact & Outcome

Gave ops and support teams a single place to spot stuck orders, shipment problems, and support trends without switching tools or asking engineering for a one-off query. Replaced a recurring stream of ad-hoc data requests with a self-serve dashboard, directly cutting down on both team context-switching and engineering interrupts.