Observability & Alerting

Structured logging, Prometheus metrics, distributed tracing, and threshold-based anomaly alerting — built into the governance pipeline.

Three Layers

GaaS observability is designed around three layers that work independently. You can run with just structured logs, add Prometheus scraping, and optionally enable OpenTelemetry tracing — each layer requires zero code changes.

1

Structured Logging

JSON-formatted structured logs via structlog. Every log entry carries a request_id for correlation across the pipeline. Set GAAS_LOG_LEVEL to control verbosity.

2

Prometheus Metrics

Scrape GET /metrics for HTTP request counts, latency histograms, pipeline stage durations, decision verdicts, error rates, circuit breaker state, and more. No authentication required on this endpoint.

3

Distributed Tracing

Optional OpenTelemetry integration. Set GAAS_TRACING_ENABLED=true and GAAS_OTEL_ENDPOINT to send spans for every pipeline stage to your collector. Zero overhead when disabled.


Anomaly Alerting

Beyond passive metrics, GaaS actively monitors the health of its own platform and raises an alert when a threshold is exceeded. The anomaly monitor evaluates metrics on every governance decision and through a periodic background check (every 60 seconds), and reports each alert to the GaaS operations team. These are platform-health alerts, not alerts about one organization: your organization is notified by email, automatically, when one of its actions is blocked or an escalation is waiting for review.

Warm-up period. The anomaly monitor uses ephemeral rolling windows for detection. Allow approximately 5 minutes after startup for accurate anomaly detection as the windows populate with baseline data.

Event Types

Event Trigger Severity
circuit_breaker_tripped Learning calibration circuit breaker freezes due to quality degradation Critical
high_error_rate More than 10 application errors in a 5-minute window Critical
block_rate_anomaly Block rate exceeds 50% over recent decisions (min 10 sample size) Warning
pipeline_degraded Any single pipeline execution exceeds 5,000ms Warning
rate_limit_spike More than 50 rate-limit rejections in a 5-minute window Warning
escalation_queue_growth Pending escalation count ≥ 10 and growing (checked every 60s) Warning

Metrics History

In addition to real-time Prometheus metrics, your organization's governance metrics are available for each past hour, computed from its own decisions and escalations:

GET /v1/dashboard/metrics/history?hours=24&limit=100

Returns one entry for each UTC hour that had decisions or escalations, oldest first: decision counts by verdict, block rate, average risk, deliberation rate, latency percentiles (p50, p99), escalations created, escalations awaiting review at the end of the hour, and the override and false-block rates over the 7 days up to it. Shadow-mode decisions are not counted.

The window is the last 30 days unless you set it: hours (1–8760) counts back from end, or give start and end as ISO 8601 date-times. limit (default 720) and offset page through the hours. Only your organization's data is included.


Prometheus Metrics Reference

Metric Type Labels
gaas_http_requests_total Counter method, path_template, status_code
gaas_http_request_duration_seconds Histogram method, path_template
gaas_active_requests Gauge —
gaas_pipeline_stage_duration_seconds Histogram stage
gaas_pipeline_total_duration_seconds Histogram —
gaas_decisions_total Counter verdict, pipeline_mode
gaas_errors_total Counter error_code
gaas_rate_limit_hits_total Counter scope
gaas_circuit_breaker_state Gauge breaker_name (0=normal, 1=frozen)
gaas_escalations_created_total Counter —
gaas_deliberation_triggered_total Counter —

Environment Variables

Variable Default Description
GAAS_LOG_LEVEL info Structlog level (debug, info, warning, error)
GAAS_TRACING_ENABLED false Enable OpenTelemetry tracing
GAAS_OTEL_ENDPOINT localhost:4317 OTLP gRPC collector endpoint
GAAS_PAGERDUTY_ROUTING_KEY — PagerDuty Events API v2 routing key for sub-processor change notices

Related Pages