Files
Stirling-PDF/engine/src/stirling/api/agent_capabilities.py
T
Anthony StirlingandGitHub 3ecd95b779 Add MCP server with OAuth/API-key auth (#6570)
Adds an optional MCP server (proprietary module) that exposes Stirling's
PDF operations and AI capabilities to MCP clients. Off by default, zero
footprint when disabled.

### What
- New `/mcp` endpoint: streamable-HTTP + JSON-RPC 2.0; 8 tools
(describe_operation, pages/convert/misc/security category tools, AI,
upload, download).
- Runs real operations over an internal loopback; results returned
inline as base64 (small) or by fileId (large).

### Auth (two modes)
- OAuth2 resource server: RFC 9728 protected-resource metadata, RFC 8707
audience binding, JWKS, `mcp.tools.read/write` scopes; binds each token
to a provisioned Stirling account.
- API-key mode: reuses Stirling per-user `X-API-KEY` (no IdP needed).

### Security
- Per-user file ownership in FileStorage: async/queued writes scoped to
the submitting user; legacy/owner-less files stay readable.
- Admin allow/block list controls which operations are exposed.
- Python engine gated behind a shared secret (`X-Engine-Auth`).
- MCP filter chain is isolated and cannot weaken the main app's
security.
- Hardened: no upstream error-body leakage, log injection sanitized,
fileId path/sidecar enumeration blocked.

### Config / footprint
- Disabled by default (`mcp.enabled=false`); all beans
`@ConditionalOnProperty`.
---

## Checklist

### General

- [ ] I have read the [Contribution
Guidelines](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/CONTRIBUTING.md)
- [ ] I have read the [Stirling-PDF Developer
Guide](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/DeveloperGuide.md)
(if applicable)
- [ ] I have read the [How to add new languages to
Stirling-PDF](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/devGuide/HowToAddNewLanguage.md)
(if applicable)
- [ ] I have performed a self-review of my own code
- [ ] My changes generate no new warnings

### Documentation

- [ ] I have updated relevant docs on [Stirling-PDF's doc
repo](https://github.com/Stirling-Tools/Stirling-Tools.github.io/blob/main/docs/)
(if functionality has heavily changed)
- [ ] I have read the section [Add New Translation
Tags](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/devGuide/HowToAddNewLanguage.md#add-new-translation-tags)
(for new translation tags only)

### Translations (if applicable)

- [ ] I ran
[`scripts/counter_translation.py`](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/docs/counter_translation.md)

### UI Changes (if applicable)

- [ ] Screenshots or videos demonstrating the UI changes are attached
(e.g., as comments or direct attachments in the PR)

### Testing (if applicable)

- [ ] I have run `task check` to verify linters, typechecks, and tests
pass
- [ ] I have tested my changes locally. Refer to the [Testing
Guide](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/DeveloperGuide.md#7-testing)
for more details.
2026-06-10 09:46:25 +00:00

161 lines
5.7 KiB
Python

"""
Curated registry of agent capabilities the MCP server (Java side) is allowed to publish.
Internal sub-agents (currently only ``ExecutionPlanningAgent`` - it lives behind the orchestrator
and has no end-user-facing API surface) are intentionally absent. The handoff spec calls for
"user-facing" capabilities only; revisit this list when adding a new agent and ask whether MCP
clients should be able to invoke it directly.
The Java side pulls ``/api/v1/agents/capabilities`` once at boot and again every few minutes; the
manifest is the authoritative source for the ``stirling_ai`` MCP tool's operation enum.
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any
from pydantic import BaseModel
from stirling.contracts import (
AgentDraftRequest,
AgentExecutionRequest,
AgentRevisionRequest,
Evidence,
FolioManifest,
PdfCommentRequest,
PdfEditRequest,
PdfQuestionRequest,
)
@dataclass(frozen=True)
class AgentCapability:
"""One row in the curated manifest.
Attributes:
id: stable capability identifier (used as the operation enum value in
``stirling_ai``). Avoid renaming - clients persist these.
description: one-line human-friendly summary shown inside MCP tool descriptions.
input_model: Pydantic class whose JSON Schema becomes the capability's
``input_schema``. Auto-derived; do not hand-write schemas.
mode: ``"sync"`` if the capability returns content inline, ``"async"`` if it returns a
plan that Java executes via the job pipeline.
required_scope: coarse OAuth scope. ``mcp.tools.read`` for pure-read capabilities
(Q&A, audits) and ``mcp.tools.write`` for anything that yields a plan / file.
route: HTTP path Java POSTs to when invoking this capability. When a capability does
not have a stable per-agent route yet, use the generic invoke fallback at
``/api/v1/agents/invoke/{id}``.
"""
id: str
description: str
input_model: type[BaseModel]
mode: str
required_scope: str
route: str
EXPOSED_CAPABILITIES: list[AgentCapability] = [
AgentCapability(
id="pdf-question-answer",
description="Answer a natural-language question about a PDF document.",
input_model=PdfQuestionRequest,
mode="sync",
required_scope="mcp.tools.read",
route="/api/v1/pdf-question",
),
AgentCapability(
id="pdf-edit-plan",
description=(
"Produce an edit plan (a structured sequence of PDF operations) from a"
" natural-language edit request. The plan is executed by Java through the job"
" pipeline; this capability does not modify files itself."
),
input_model=PdfEditRequest,
mode="async",
required_scope="mcp.tools.write",
route="/api/v1/pdf-edit",
),
AgentCapability(
id="agent-draft",
description=(
"Draft a structured agent specification from a free-text description of the task the user wants automated."
),
input_model=AgentDraftRequest,
mode="sync",
required_scope="mcp.tools.read",
route="/api/v1/ai/agents/draft",
),
AgentCapability(
id="agent-revise",
description=("Revise an existing draft agent specification based on user feedback or constraint changes."),
input_model=AgentRevisionRequest,
mode="sync",
required_scope="mcp.tools.read",
route="/api/v1/ai/agents/revise",
),
AgentCapability(
id="math-audit-examine",
description=(
"Examine a folio manifest of financial / numeric documents and surface the"
" evidence that needs to be checked for arithmetic consistency."
),
input_model=FolioManifest,
mode="sync",
required_scope="mcp.tools.read",
route="/api/v1/ai/math-auditor-agent/examine",
),
AgentCapability(
id="math-audit-deliberate",
description=(
"Render a deliberated verdict on a single piece of evidence the examine step"
" surfaced (does the arithmetic check out, with what caveats)."
),
input_model=Evidence,
mode="sync",
required_scope="mcp.tools.read",
route="/api/v1/ai/math-auditor-agent/deliberate",
),
AgentCapability(
id="pdf-comment-generate",
description="Generate inline review comments for a PDF document.",
input_model=PdfCommentRequest,
mode="sync",
required_scope="mcp.tools.read",
route="/api/v1/pdf-comment/generate",
),
AgentCapability(
id="agent-next-action",
description=(
"Decide the next execution step for an in-progress agent workflow. Returns a"
" ToolCall, Completed, or CannotContinue action."
),
input_model=AgentExecutionRequest,
mode="sync",
required_scope="mcp.tools.read",
route="/api/v1/agents/next-action",
),
]
def manifest_payload() -> dict[str, Any]:
"""Serialize the curated registry to the wire shape consumed by Java.
Schema is derived from ``input_model.model_json_schema()`` so we never hand-write JSON
Schema - the Pydantic model is the single source of truth.
"""
items: list[dict[str, Any]] = []
for cap in EXPOSED_CAPABILITIES:
items.append(
{
"id": cap.id,
"description": cap.description,
"input_schema": cap.input_model.model_json_schema(),
"mode": cap.mode,
"required_scope": cap.required_scope,
"route": cap.route,
}
)
return {"version": 1, "capabilities": items}