AIOps Observability Platform

for Data Pipelines

Databricks monitoring, run review, and alerts for failures, cancellations, and telemetry problems are implemented. Prometheus scrapes the exporter and stores metrics and run checkpoints; Loki stores captured run messages. PostgreSQL holds the registry and alert history. Optional Teams and webhook notifications, expiring mutes, and AI analysis are available. Full log analysis, persistent investigations, automated recovery, lineage, and prediction remain planned.

Key capabilities

1. Centralized observability

  • Unified monitoring for registered data pipelines.
  • Clear workload, collection, and telemetry health without direct provider reads from the dashboard.
  • Execution duration, output, freshness, throughput, SLA, lineage, and dependency evidence where each signal is actually available.
  • A useful health model for every monitored resource.

Status: Core monitoring path implemented; broader observability pending.

What I have done

  • Admins connect Databricks and register jobs or pipelines. Credentials stay encrypted; provider validation and collection have separate entrypoints.
  • Prometheus scrapes registered resources. The exporter has no database dependency or collection schedule; dashboard reads use stored telemetry.
  • The dashboard separates workload health, collection failures, stale telemetry, and query outages. Search and filters cover the full fleet.
  • Run review reads paginated checkpoints from Prometheus and captured messages, native state, and provider links from Loki. Missing context and store outages are shown explicitly.
  • The single-host Compose stack includes web, exporter, PostgreSQL, Temporal, Prometheus, Loki, and Alertmanager.

What remains

  • Verify fleet pagination, filters, provider collection, and failure recovery in the target deployment.
  • Add full execution and infrastructure logs. Captured provider messages are not a complete log archive.
  • Add lineage, dependency views, numeric health scoring, table freshness, throughput, and SLA evaluation. Telemetry freshness does not measure table freshness.
Implementation evidence
  • packages/contracts/src/read.ts
  • packages/db/src/monitoring-registry.ts
  • apps/databricks-exporter/src/probe.ts
  • apps/databricks-exporter/src/metrics.ts
  • apps/web/src/server/services/prometheus/monitoring-read.service.ts
  • apps/web/src/server/trpc/router/monitoring.ts
  • apps/web/src/ui/routes/_authenticated/index.tsx
  • apps/web/src/ui/routes/_authenticated/pipelines/$pipelineId.tsx
  • apps/web/src/ui/data/monitoring-range.ts
  • apps/web/src/ui/components/pipelines/pipeline-resource-details.tsx
  • apps/web/src/server/trpc/router/executions.ts
  • apps/web/src/ui/components/pipelines/pipeline-execution-history.tsx
  • docker/compose.deploy.yml
  • docs/fleet-pipeline-data-flow.md
  • apps/web/src/server/services/runs/prometheus-runs.service.ts
  • apps/web/src/server/services/runs/run-context.service.ts
  • apps/web/src/server/services/runs/run-reads.isolation.test.ts

2. Intelligent failure detection

  • Detect failed, delayed, or anomalous pipeline executions.
  • Correlate failures across workflows, jobs, notebooks, clusters, infrastructure, and dependencies.
  • Detect recurring patterns and warn before SLA breaches.

Status: Persistent built-in detection implemented; correlation and prediction pending.

What I have done

  • Normalize provider outcomes into queued, running, succeeded, failed, canceled, skipped, and unknown states. Failed runs default to critical and canceled runs to warning; configured severity is shared across dashboard health, alerts, and notifications.
  • Keep workload health separate from collection health, telemetry freshness, and Prometheus availability. Retained observations keep their original timestamps.
  • Prometheus evaluates run-failure, run-cancellation, collection-failure, and stale-telemetry rules with shared defaults and per-resource overrides. Cancellation alerts count canceled attempts since the latest success.
  • Persist alert history across reconciliation retries and restarts. Missing telemetry or a running retry does not prove recovery; disabling monitoring records a separate closure reason.

What remains

  • Add cross-resource correlation, recurring-pattern detection, anomaly models, and predictive SLA warnings. Overdue-run detection is outside the current release.
  • Validate detector quality against real incidents. Duration and outcome summaries do not establish anomalies.
Implementation evidence
  • packages/contracts/src/read.ts
  • apps/databricks-exporter/src/metrics.ts
  • apps/web/src/server/services/prometheus/monitoring-read.service.ts
  • apps/web/src/ui/data/pipelines.ts
  • apps/web/src/ui/data/monitoring-notices.ts
  • apps/web/src/ui/routes/_authenticated/issues.tsx
  • docs/alerting-operations.md
  • packages/db/src/alerting-store.ts
  • apps/web/src/server/services/alerting/reconciler.ts
  • apps/web/src/server/services/alerting/reconciler.test.ts
  • packages/contracts/src/health.ts
  • packages/contracts/src/health.test.ts
  • packages/contracts/src/alert.ts

3. AI-powered log analysis

  • Collect execution and supporting-infrastructure logs.
  • Identify probable causes and summarize technical evidence in plain language.
  • Categorize incidents across authentication, data quality, schema, compute, and dependency failures.
  • Use historical incidents and knowledge sources to recommend likely resolutions.

Status: Monitoring and captured-error analysis implemented; full log analysis pending.

Analysis available

  • Optional AI analysis combines monitoring metrics with the current failed run's captured error message. Older failures are excluded after a running or successful outcome.
  • Return a short summary, explanation, impact, recommendation, and confidence label, identifying AI output or deterministic fallback.
  • Bound generation, deduplicate requests, cache by evidence, and fall back when the model is unavailable. Cause and downstream impact remain unknown unless supported by evidence.
  • Run review displays bounded, redacted provider messages stored in Loki. Dashboard and AI reads never contact Databricks.

Log analysis still required

  • Implement full execution and infrastructure log ingestion, retention, access control, redaction, and run correlation. Provider log collection remains disabled.
  • Add evidence-backed incident categories and retrieval of historical resolutions and knowledge sources.
  • Evaluate diagnosis quality, grounding, redaction, and usefulness against representative incidents.
Implementation evidence
  • apps/web/src/server/services/ai/pipeline-analysis.service.ts
  • apps/web/src/server/services/ai/pipeline-analysis.service.test.ts
  • apps/web/src/ui/components/pipelines/pipeline-ai-summary.tsx
  • apps/web/src/ui/components/pipelines/pipeline-run-review.tsx
  • apps/databricks-exporter/src/index.ts
  • apps/databricks-exporter/src/run-context.ts
  • apps/web/src/server/services/runs/current-failure-context.service.ts
  • apps/web/src/server/services/runs/current-failure-context.service.test.ts

4. AI recommendation engine

For each incident, the platform should provide:

  • A grounded cause summary and confidence signal.
  • Actionable remediation with the expected operational or business impact.
  • Preventive measures and related successful historical resolutions.

Status: Bounded monitoring recommendations implemented.

What I have done

  • Surface structured summary, explanation, impact, recommendation, and high, medium, or low confidence fields in monitoring views.
  • Label whether analysis came from the configured AI model or deterministic fallback logic.
  • Base recommendations on monitoring state and the current failed run's captured error message; leave unsupported causes and downstream impact unknown.
  • State when cause and impact are not established instead of presenting metric correlation as diagnosis.

What remains

  • Add calibrated confidence if a numeric score is required; the current labels are not probabilities.
  • Use authorized logs, lineage, business criticality, and historical outcomes as explicit evidence sources before making richer cause or impact claims.
  • Add multi-step remediation playbooks, preventive measures, related successful resolutions, and outcome feedback.
  • Keep recommendations separate from execution. No provider retry, repair, or automated remediation action is implemented.
Implementation evidence
  • packages/contracts/src/analysis.ts
  • apps/web/src/server/services/ai/pipeline-analysis.service.ts
  • apps/web/src/ui/components/pipelines/pipeline-ai-summary.tsx
  • apps/web/src/ui/components/pipelines/pipeline-issues.tsx

5. Human-in-the-loop recovery

Operators should be able to review evidence and then:

  • Approve or reject an AI diagnosis.
  • Trigger an audited remediation or rerun.
  • Execute failed-task or checkpoint recovery when the provider supports it.
  • Escalate incidents for manual intervention.

Status: Local review workflow available; persistence and recovery pending.

Interface available

  • Built an Issues board with New, Investigating, Waiting for approval, and Resolved columns.
  • Keep cards from the latest monitoring refresh visible while the page is open and allow operators to move them between columns.
  • Link issue cards to resource detail pages and show current monitoring analysis and detections.
  • Implemented admin controls to enable, disable, and remove monitored resources.

Workflow and actions required

  • Implement issue #11 persistence for stages, assignments, approvals, stable evidence links, permissions, concurrent updates, and audit history. Board moves currently live only in React state and disappear on refresh.
  • Expand run review beyond captured provider messages to authorized execution and infrastructure logs. Run links already select stored checkpoints and context.
  • Implement diagnosis approval and rejection, escalation, provider rerun or failed-task repair, checkpoint restart, and automated remediation as audited actions.
  • Capture operator feedback and recovery outcomes for later recommendation evaluation.
Implementation evidence
  • apps/web/src/ui/routes/_authenticated/issues.tsx
  • apps/web/src/ui/data/issue-board.ts
  • apps/web/src/ui/components/pipelines/pipeline-execution-history.tsx
  • apps/web/src/ui/components/pipelines/pipeline-run-review.tsx
  • apps/web/src/ui/routes/_authenticated/pipelines/$pipelineId.tsx
  • apps/web/src/server/trpc/router/resources.ts

6. Automated incident management

  • Create and retain incidents for critical failures.
  • Notify configured external destinations without making them required for in-app alerts.
  • Prioritize incidents using business criticality and dependencies.
  • Measure detection and resolution performance.

Status: Persistent alerts, notifications, and mutes implemented; incident workflows pending.

Alerting release implemented

  • Prometheus evaluates failure, cancellation, collection, and stale-telemetry rules; PostgreSQL retains active and resolved alerts, including recovered and monitoring-disabled closure reasons.
  • A Temporal worker reconciles alert history, publishes configuration, and retries failed synchronization.
  • Provide fleet and resource alert views, shared defaults, and per-resource overrides. Completed alert history has a configurable 30-day retention default.
  • Admins configure encrypted Teams Workflows or generic webhook destinations and send test alerts. Alertmanager handles grouping, retries, and firing notifications.
  • Admins and members can mute a resource alert until a chosen expiry or cancel its mutes. Pending and failed synchronization remain visible and retryable.
  • Admins can mute warning notifications across all destinations. Warning health and history remain visible; critical notifications and explicit destination tests stay enabled.
  • Notifications are scoped to active monitoring periods. In-app history remains available when external delivery fails.

Broader incident management remains

  • Verify Teams delivery and routing in the target environment. Automated receiver tests do not prove that a real Teams channel received a card; applied configuration does not confirm delivery.
  • Add business-criticality prioritization, escalation, and detection and resolution time measurements.
  • Implement persistent incident assignment, approval, and audit history. The separate Issues board remains local to the browser.
Implementation evidence
  • apps/web/src/ui/data/pipelines.ts
  • apps/web/src/ui/routes/_authenticated/issues.tsx
  • apps/databricks-exporter/src/probe.ts
  • apps/databricks-exporter/src/metrics.ts
  • packages/db/src/alerting-store.ts
  • apps/web/src/server/services/alerting/engine/rules.ts
  • apps/web/src/server/services/alerting/reconciler.ts
  • apps/web/src/ui/routes/_authenticated/alerts.tsx
  • apps/web/src/ui/components/alerts/alert-settings.tsx
  • apps/web/src/server/temporal/workflows.ts
  • docs/alerting-operations.md
  • apps/web/src/server/services/alerting/engine/alertmanager-config.ts
  • apps/web/src/server/services/alerting/sync-mute.ts
  • apps/web/src/server/services/alerting/engine/teams-delivery.integration.bun.test.ts
  • docs/teams-alerting.md
  • apps/web/src/ui/components/alerts/alert-settings.test.tsx

7. Predictive operations

Predictive operations should:

  • Forecast pipeline failures and SLA violations.
  • Detect capacity bottlenecks and recommend scaling.
  • Identify pipelines that need optimization.

Status: Telemetry foundation available; prediction pending.

Telemetry foundation

  • Collect bounded recent execution outcomes, latest duration, latest output rows when available, and collection-health evidence.
  • Store duration samples in Prometheus and display a gap-aware duration history for fresh telemetry.
  • Provide recent outcome counts without presenting them as a long-term reliability score.
  • Removed the previous synthesized expected-volume baseline because no baseline model or evidence supported it.

Predictive work required

  • Define prediction targets, minimum history, features, labels, baselines, evaluation windows, and acceptance thresholds.
  • Collect the longer-term workload, schedule, dependency, and compute evidence required for failure, SLA-risk, and capacity models.
  • Implement and validate failure forecasting, bottleneck detection, scaling recommendations, and optimization ranking.
  • Distinguish a measured prediction from a current-state rule, duration comparison, or AI-written summary.
Implementation evidence
  • apps/databricks-exporter/src/metrics.ts
  • packages/contracts/src/read.ts
  • apps/web/src/server/services/prometheus/monitoring-read.service.ts
  • apps/web/src/ui/components/pipelines/duration-history-chart.tsx
  • docs/metrics.md