Plumbline

LLM answers checked against the data, or a refusal with the reason.

Connect your own LLM app to a dataset through Plumbline. Strict answers publish figures only after checking them against SQL facts; broader or unsupported questions are labeled or refused with the reason, never passed off as verified. No account: the demo endpoint is open, rate-limited per address.

Terms used here

MCP
Model Context Protocol: the standard by which LLM apps (Claude, ChatGPT, VS Code, others) call external tools. Plumbline is an MCP server; you add its URL to your LLM app.
snapshot
A pinned copy of the dataset that answers come from. The demo snapshot is a 100-person, 19-department payroll dataset from an HR sandbox (synthetic people, realistic structure).
verified
Each published figure in the strict answer path is matched against SQL-computed facts from the snapshot. Figures that match are marked verified; a mismatch is flagged as a conflict in front of you. Other tiers stay explicitly labeled.
refusal
When the snapshot has no evidence for a question, the answer is a refusal naming the gap. A refusal is a correct outcome, not an error.
explanation record
The audit trail behind an answer: which facts were used, what was checked, what was refused or disclosed. Retrievable per answer.

What it looks like in your LLM app

mockup · your LLM app
Claude · Plumbline connected over MCP
What will compensation look like next year?
PLUMBLINE
refused on the question, before any drafting prompt was issued
Who are our biggest pay outliers versus their own department, and by how much?
PLUMBLINE

Five people sit well above their department's median pay. Florian Grady leads at 6.0× the Electronics median, $518,878 against $86,590.

up to department medianabove median (the outlier gap)
5/5 figures store-verified · sources: compensation_facts, department_growth

Mockup of an LLM chat with Plumbline connected: you type the question, the verdict panel renders in the conversation. Every figure and name shown is real output against the demo snapshot (annualized pay). Hosts without panel support get the same content as text.

What you are connecting to

host layerThe panel and gateway

Serves MCP servers into Claude, ChatGPT, VS Code, and other MCP hosts: the in-chat panel, confirm gates, and a plain-text fallback where panels are unavailable.

trust pipelinePlumbline

Every figure in an answer is checked against SQL-computed facts and labeled with how it was checked; a question the evidence cannot support is refused with the reason. Workforce data is the first domain; the mechanics are domain-agnostic.

Connect your LLM

A shared demo endpoint is open. No account; it is rate-limited per address and per day. Pick your client:

terminal, then restart the session

claude mcp add --transport http plumbline-demo https://plumbline.lattice-sys.com/mcp/demo

Settings → Connectors → Add custom connector → paste as URL

https://plumbline.lattice-sys.com/mcp/demo

Desktop then runs its connect step and approves the demo automatically (no account, no password). The URL must be publicly reachable over https, which is what this address is.

Settings → Connectors → Advanced → Developer mode → Add connector → MCP server URL

https://plumbline.lattice-sys.com/mcp/demo

ChatGPT connects from its own servers, so the URL must be https and publicly reachable; a private/tailnet address will not work from ChatGPT.

.vscode/mcp.json in your workspace

{ "servers": { "plumbline-demo": { "type": "http", "url": "https://plumbline.lattice-sys.com/mcp/demo" } } }

~/.cursor/mcp.json

{ "mcpServers": { "plumbline-demo": { "url": "https://plumbline.lattice-sys.com/mcp/demo" } } }

~/.gemini/settings.json

{ "mcpServers": { "plumbline-demo": { "httpUrl": "https://plumbline.lattice-sys.com/mcp/demo" } } }

any MCP client speaking streamable HTTP

{"url": "https://plumbline.lattice-sys.com/mcp/demo"}

Then type this in your LLM session

Paste these prompts one at a time. Each shows a different part of the pipeline. In hosts with MCP panel support (Claude, ChatGPT) the verdict renders as an in-chat panel with suggestion chips; a chip copies its question so you can paste it here.

Ask one subject per question. Facts are computed for a single family, so a question spanning two subjects gets a facts block for one of them and leaves the other half unanswered.

1 · a question the data cannot answer

Using the plumbline tools, answer exactly this question: What will compensation look like next year?

Expected: a refusal naming the gap (this snapshot records what has already happened and holds no forecast), returned before the model is given anything to draft from, so the refusal does not depend on the model declining to guess.

2 · a question it can answer

Using the plumbline tools, answer exactly this question: How many people and departments are in this company?

Expected: a verified answer (100 people, 19 departments) with per-figure provenance.

3 · analysis with per-figure provenance

Using the plumbline tools: which roles are unusually paid, and which employees are compensation outliers?

Expected: named outliers with each figure carrying its own grade (verified, corroborated, derived, or no verified source).

4 · a refusal on policy, not on data

Using the plumbline tools, answer exactly this question: Which employee most deserves a raise this year, based on their performance?

Expected: refused as an employment decision with no performance evidence in the snapshot. The data could produce a name; the pipeline declines to.

5 · what a department costs, with its own caveats

Using the plumbline tools, answer exactly this question: What is our biggest cost center?

Expected: departments ranked by salary cost, carrying the three disclosures this family always attaches. The basis is annualized salary and not payroll, so the figures must not be reconciled against the payroll totals; every row states how many of its people the money actually covers (Movies is 11 of its 12, and company-wide 12 of 100 people have no annualizable rate); and the rows shown are the ten most expensive of the 19, so the last one listed is not the cheapest.

6 · org topology

Which managers look like routing hot spots?

Expected: grounded topology facts, plus a caveat naming the 6 of 20 managers who are marked inactive in the directory (terminated, retired, or prehire) yet still own 24 direct reports between them, so a stale reporting line is not presented as a current one.

7 · anomalies against peers

Find anomalies in benefits deduction patterns compared to peers.

Expected: anomaly findings tied to snapshot facts; partial evidence stays visible instead of becoming a guess.

8 · sandboxed code, then earn the figure back

Get the total headcount and the three largest department counts from Plumbline's verified tools, use run_code to compute each department's share of the total in the isolated sandbox (show the arithmetic), then corroborate the total with cross_check (value_key 'people').

Expected: the sandbox arithmetic comes back labeled UNVERIFIED (it bypasses the trust path by design); cross_check then corroborates the total against the SQL fact store. That round trip is the point: computed figures are cheap, corroborated figures are earned.

9 · the audit trail

Open the evidence receipt linked from that last answer.

Expected: the receipt shows which facts were used, what was checked, and what was refused or disclosed.