57dc91585d
Internal SmartGift build of a Claude Code monitoring dashboard. Lanes: a durable unit of parallel agent work, one per working directory, tracked across session restarts. Managed lanes are git worktrees the dashboard provisions and can reset or remove behind a three-check destroy guard and a counted preflight; adopted lanes are directories you already own and are never destroyable. Pipelines: a lane moves through pipeline stages. A stage the agent declares with evidence renders green; a stage inferred from the tool-event stream renders dashed amber and never counts as done. Detection is forward-only within a 30-minute window, and never writes the declared stage. Workspace: one page at /run with a lane grid, the selected lane's pipeline, and a full Claude console behind a disclosure.
Coralogix Integration
Full-stack observability for Claude Code Agent Monitor via Coralogix — logs, metrics, traces, and SLO tracking through a single platform.
Architecture
graph TB
subgraph "Kubernetes Cluster"
APP["Agent Monitor Pods"]
MCP["MCP Sidecar"]
OTEL["OTel Collector<br/>(DaemonSet)"]
end
APP -->|"metrics + logs"| OTEL
MCP -->|"metrics + logs"| OTEL
OTEL -->|"OTLP (gRPC)"| CX["Coralogix Platform"]
subgraph "Coralogix"
CX --> LOGS["Log Analytics<br/>DataPrime Queries"]
CX --> MET["Metrics<br/>PromQL + Recording Rules"]
CX --> TRACE["Distributed Tracing"]
CX --> ALERT["Alert Engine"]
CX --> DASH["Custom Dashboards"]
CX --> SLO["SLO Management"]
end
ALERT -->|"Critical"| PD["PagerDuty"]
ALERT -->|"Warning"| SLACK["Slack"]
style OTEL fill:#4f46e5,color:#fff
style CX fill:#1a1a2e,color:#fff
style LOGS fill:#7c3aed,color:#fff
style MET fill:#e6522c,color:#fff
style TRACE fill:#059669,color:#fff
style ALERT fill:#dc2626,color:#fff
style DASH fill:#f46800,color:#fff
style SLO fill:#0ea5e9,color:#fff
Files
| File | Purpose |
|---|---|
values.yaml |
Helm values for Coralogix OpenTelemetry Collector |
alerts.yaml |
Alert definitions (mirrors Prometheus/Alertmanager rules) |
dashboards.yaml |
Custom dashboard with 6 rows, 18 panels, SLO tracking |
coralogix-terraform.tf |
Terraform-managed alerts, parsing rules, recording rules |
Quick Start
1. Add the Helm Repository
helm repo add coralogix https://cgx.jfrog.io/artifactory/coralogix-charts-virtual
helm repo update
2. Create the API Key Secret
kubectl create secret generic coralogix-keys \
--namespace agent-monitor \
--from-literal=PRIVATE_KEY=<YOUR_CORALOGIX_SEND_YOUR_DATA_KEY>
3. Deploy the OTel Collector
helm install coralogix-otel coralogix/opentelemetry \
--namespace agent-monitor \
-f deployments/monitoring/coralogix/values.yaml
4. Import the Dashboard
Upload dashboards.yaml via the Coralogix UI:
Dashboards → Custom Dashboards → Import
5. (Optional) Terraform-managed Alerts
cd deployments/monitoring/coralogix
export CORALOGIX_API_KEY="<your-key>"
export CORALOGIX_ENV="coralogix.com"
terraform init
terraform apply
What Gets Collected
| Signal | Source | Destination |
|---|---|---|
| Logs | Pod stdout/stderr (JSON structured) | Coralogix Log Analytics |
| Metrics | Prometheus scrape (/api/health) |
Coralogix Metrics |
| K8s Metrics | kubelet, cAdvisor, host metrics | Coralogix Metrics |
| Traces | OTLP from application (if instrumented) | Coralogix Tracing |
Alert Parity
All 10 Prometheus/Alertmanager rules are replicated in Coralogix:
| Alert | Severity | Prometheus | Coralogix |
|---|---|---|---|
| Instance Down | Critical | ✓ | ✓ |
| High Error Rate | Critical | ✓ | ✓ |
| Pod Restart Loop | Critical | ✓ | ✓ |
| PV Nearly Full | Critical | ✓ | ✓ |
| High Latency | Warning | ✓ | ✓ |
| WebSocket Spike | Warning | ✓ | ✓ |
| High Memory | Warning | ✓ | ✓ |
| High CPU | Warning | ✓ | ✓ |
| HPA Maxed Out | Warning | ✓ | ✓ |
| Slow DB Queries | Warning | ✓ | ✓ |
Dashboard Panels
The custom dashboard provides 18 panels across 6 rows:
- Overview — Active sessions, request rate, WebSocket connections
- HTTP Performance — Latency distribution, error rate, status codes
- Application Logs — Error log stream (DataPrime), log volume by severity, hook throughput
- Infrastructure — CPU, memory, pod status
- Database & Storage — SQLite query duration, PV usage, network I/O
- SLO Tracking — Availability SLO (99.9%), latency SLO (P95 < 500ms), error budget burn
Coralogix Regions
Set global.domain in values.yaml to match your Coralogix region:
| Region | Domain |
|---|---|
| US1 | coralogix.us |
| US2 | cx498.coralogix.com |
| EU1 | coralogix.com |
| EU2 | eu2.coralogix.com |
| AP1 (India) | coralogix.in |
| AP2 (Singapore) | coralogix.sg |