Quick Start¶
Prerequisites¶
- Docker and Docker Compose
- Access to Prometheus, Loki, and Alertmanager endpoints
- Environment variables configured (see below)
1. Configure Environment¶
Copy the example env file and fill in your endpoints:
Edit scripts/.env:
# Required
PROMETHEUS_URL=https://prometheus-read.example.com/select/0/prometheus
LOKI_URL=https://loki.example.com
ALERTMANAGER_URL=https://alertmanager.example.com
# Optional
CLUSTER_NAME=my-cluster
ML_ENABLED=true
EXCLUDE_NAMESPACES_CSV=kube-system,kube-node-lease
2. Start the Stack¶
This builds Go binaries and starts:
- 1× Controller (port 8080 for metrics)
- 3× Workers (gRPC port 50052)
- 1× Redis (port 6379)
- 1× ML Service (gRPC port 50051, metrics port 8082)
3. Monitor¶
4. Stop¶
5. Verify It's Working¶
After starting, check:
# Health check
curl -s http://localhost:8080/readyz
# Expected: 200 OK (or 503 if a dependency is unreachable)
# Metrics
curl -s http://localhost:8080/metrics | grep staffops_ad_controller_cycle
# Expected: cycle_duration_seconds histogram with increasing count
Healthy indicators
staffops_ad_controller_cycle_duration_seconds_countincreasing every 30sstaffops_ad_worker_queries_totalincreasingstaffops_ad_detection_anomalies_total> 0 (some anomalies expected)- Logs show
[DRY-RUN] would fire alertmessages
Common issues
- 503 on /readyz: Check that PROMETHEUS_URL, LOKI_URL, ALERTMANAGER_URL are reachable from inside Docker
- No anomalies: Wait 30+ minutes for baselines to warm up (60 samples × 30s)
- ML errors: Check ML service logs:
docker compose -f scripts/docker-compose.yaml logs ml