Skip to content

AIgent Squad Chat — Architecture

Option A: thin client behind the Grafana LLM app

Browser (this plugin, React)                 — no secret, no endpoint in the browser
  │  @grafana/llm streamChatCompletions()      (streams over Grafana Live)
  ▼
Grafana LLM app (grafana-llm-app)             — holds the AIgent Squad URL + token
  │  OpenAI-compatible /v1/chat/completions
  ▼
AIgent Squad (external)                       — subagents + tools (Loki/Tempo/PromQL/K8s)

AIgent Squad is registered as a custom OpenAI-compatible provider inside the Grafana LLM app (provider type openai with a custom url). The plugin never sees the endpoint or token — they live in the LLM app's secureJsonData.

Why Option A (vs a Go backend)

Option A — LLM app Option B — Go backend in the plugin
Secret handling LLM app holds it server-side Plugin backend holds it
New moving parts Reuses an installed app Adds a Go backend to a frontend-only repo
Streaming @grafana/llm over Grafana Live Custom SSE passthrough
Chosen ✅ Fallback only

Option A was validated end-to-end (Grafana 13.0.2 + grafana-llm-app v1.0.8): Bearer auth forwarded, streaming works over Grafana Live, and long tool-call pauses survive without timeout.

Provider URL gotcha

grafana-llm-app appends /v1/ to the configured base URL. Configure the provider URL without the /v1 suffix (e.g. https://aigent-squad.<domain>), otherwise requests hit /v1/v1/chat/completions and 404.

Guardrail compatibility — no system message

AIgent Squad's Bedrock guardrail treats an injected system prompt as untrusted input and blocks it (HTTP 403). The plugin therefore sends no system message — the squad has its own system prompt. Instead, the latest user turn is prefixed with a configurable source tag ([source:grafana-plugin/aigent-squad-chat]) for audit/filtering on the squad side.

Root fix on the squad side

The squad also strips system messages before the guardrail (so any OpenAI client — LibreChat, Continue.dev — works). The plugin's no-system-message behavior is the client-side half of that contract.

Streaming behavior

AIgent Squad pseudo-streams: it runs to completion (executing tools), then emits the answer as a burst of SSE chunks. During the wait the UI shows the thinking indicator; when the burst arrives, @grafana/llm's accumulateContent() renders it progressively. True token-by-token streaming is deferred to a future squad release.

Frontend safety timers:

Timer Default Purpose
Thinking-idle 3s Switch to the "working…" indicator when no chunk arrives.
Max stream idle 120s Hard error if no chunk of any kind arrives (tool runs can be long).

Both are adjustable on the Configuration page.