Skip to main content
Milford is one Node process with no database. The image is built for linux/amd64 and linux/arm64, so it runs on a Raspberry Pi or a small server.

Docker

The HTTP server image is ghcr.io/milfordai/milford, and the MCP server image is ghcr.io/milfordai/milford-mcp. Mount a directory that holds milford.config.yaml at /config. The image reads MILFORD_CONFIG, which defaults to /config/milford.config.yaml.
Flow file paths in the config are relative to the config file, so ./flows/triage.json resolves to /config/flows/triage.json.

Docker Compose

The repository’s docker-compose.yml builds the images from source. Put milford.config.yaml and a flows/ directory next to docker-compose.yml, then set your keys:
The compose file mounts ./milford.config.yaml and ./flows read-only under /config and passes these variables through: MILFORD_TOKEN, ANTHROPIC_API_KEY, OPENAI_API_KEY, GROQ_API_KEY, TYPESAFE_API_KEY. Add the variables your config uses to the environment list.
The image healthcheck calls /health on port 8080. If you change server.port, update the healthcheck.

MCP server image

The Dockerfile has a second target for the MCP server. Set mcp.transport: http and mcp.auth.tokens in the config, then run it with the mcp profile:
It listens on port 8090 and its healthcheck calls /health there. The HTTP server image stays the default target.

Image publishing

The docker GitHub Actions workflow builds multi-arch images when a v* tag is pushed. It pushes the HTTP server to ghcr.io/milfordai/milford and the MCP server to ghcr.io/milfordai/milford-mcp, each tagged with the version and latest.

Running offline

Milford works without internet when its providers are local: a classifier behind the http provider, or a local OpenAI-compatible server. Use fallback to prefer local providers and fall back to a cloud provider.

Without Docker

Install the server from npm and run it:
The MCP server is @milfordai/mcp. See installation for global installs and the library packages.

Reliability settings

Set these before you expose the server to other clients:
  • Keep server.auth.tokens set. Without tokens the API is open.
  • Tune run.timeoutMs, run.maxConcurrentRuns and server.maxBodyBytes to your workload. The defaults are 60 seconds, 64 runs and 1 MB.
  • Add circuitBreaker and a fallback to providers that can go down, and rateLimit to providers with quotas. See providers.