Quick Start¶
Get from zero to serving predictions in five minutes.
Prerequisites¶
- Docker with Compose v2
uv≥ 0.4 and Python 3.12+
Install uv if needed:
0. Just the CLI¶
The full stack below is what you want on a workstation. If all you need is to talk to an ExaMLOps platform that already runs somewhere — a login node, a laptop, a CI job — install the CLI on its own. It has no Docker, no MLflow and no GPU dependency:
uv pip install examlops # or: pip install examlops
exa --version
exa status # points at the URLs in your active context
The package reaches PyPI with the first release published through the release workflow
(Releases). Until then, install the same wheel from a checkout:
uv pip install ./platform/cli.
Heavier capabilities are extras, so the base install stays small enough for a login node. Each one is lazily imported, and a command that needs a missing extra says which to install rather than failing with a traceback:
| Extra | What it adds |
|---|---|
examlops[analysis] |
Statistical A/B analysis — exa serve ab analyze |
examlops[backup] |
Object-store and off-site backup tiers — exa backup |
examlops[coordination] |
Redis-backed cross-host locks, rate limits, deduplication, and event streams |
examlops[fairness] |
Fairlearn's MetricFrame for fairness slice metrics — exa fairness (a pure-Python fallback gives the same numbers without it) |
examlops[finops] |
YAML/expression calculation providers — user-authored cost and carbon formulas |
examlops[mcp] |
Serve the platform to LLM agents — exa mcp serve |
examlops[oidc] |
Validate OIDC access tokens (RS256 against a JWKS) |
examlops[postgres] |
Talk to a Postgres datastore instead of SQLite |
examlops[serving-sglang] |
In-process SGLang engine (GPU host) |
examlops[serving-vllm] |
In-process vLLM engine for offline batch scoring (GPU host) |
examlops[synth] |
Synthetic data generation and its release gate — exa data synth |
examlops[vector] |
pgvector vector store (EXAMLOPS_VECTOR_BACKEND=pgvector), independent of where platform state lives |
Combine them as usual — uv pip install 'examlops[mcp,analysis]'. [dev] on the repository
root pulls the whole development environment including the test dependencies.
1. Start the full stack¶
This starts Postgres, MLflow, Prefect, and Ray Serve in Docker, then installs all Python dependencies. Once complete you should see:
| Service | URL |
|---|---|
| MLflow UI | http://localhost:15000 |
| Prefect UI | http://localhost:14200 |
| Ray Serve API | http://localhost:18001 |
| Ray Dashboard | http://localhost:18265 |
| ExaMLOps Dashboard | http://localhost:18099 |
| JupyterHub | http://localhost:18888 |
2. Train a model¶
Run all discovered models with dummy data (no downloads, completes in seconds):
This runs the full pipeline for every registered model × dataset pair:
data extraction → HPC submit → evaluate → MLflow log → promote
To train a specific model:
3. Promote to Production¶
The pipeline promotes automatically if RMSE is below the threshold defined in the model config. Check the MLflow UI to confirm the @Production alias was set:
4. Serve a prediction¶
Phase 3 keeps Ray Serve auto-synced with MLflow via polling + Prefect webhook, so you usually don't need a manual reload. If you want to force one:
Send a prediction. Phase 3: optionally pin a specific lifecycle alias or raw version per request.
# Default (uses the @Production alias):
curl -X POST http://localhost:18001/predict/JPCP \
-H "Content-Type: application/json" \
-d '{"features": {"feature_0": 1.2, "feature_1": 0.8, "feature_2": 3.4}}'
# Hit the Canary alias instead:
curl -X POST http://localhost:18001/predict/JPCP \
-H "Content-Type: application/json" \
-d '{"features": {"feature_0": 1.2}, "alias": "Canary"}'
# Hit a specific historical version (lazy-loaded via the LRU cache):
curl -X POST http://localhost:18001/predict/JPCP \
-H "Content-Type: application/json" \
-d '{"features": {"feature_0": 1.2}, "version": "3"}'
Response (the alias and model_version echo what was actually served):
{
"model_name": "JPCP",
"alias": "Production",
"model_version": "1",
"run_id": "abc123...",
"prediction": 142.7
}
5. Trigger a retrain via the control plane (Phase 4)¶
Before retrain works, a Prefect deployment must exist. Run this once in a separate terminal (it registers the deployment and runs a worker — keep it running):
Then trigger a retrain from your original terminal:
Or via the exa CLI:
Or directly with curl:
curl -X POST http://localhost:18002/retrain \
-H "Authorization: Bearer ${CONTROL_PLANE_TOKEN:-changeme}" \
-H "Content-Type: application/json" \
-d '{"model_name":"JPCP","dataset_name":"PM100Dataset","is_dummy":true}'
See Control Plane for the full API.
6. Try a different data source (Phase 1)¶
# Pull from MinIO (after seeding s3://examlops-data/PM100/job_table.parquet):
exa pipeline run --model JPCP --dataset PM100Dataset --backend minio
# Snapshot the dataplane simulator and train on it:
exa pipeline run --model JPCP --dataset PM100Dataset --backend dataplane --dummy
7. Scaffold a new model (Phase 2)¶
exa scaffold DemoAD --task anomaly_detection --type classification
exa pipeline list # confirm DemoAD appears
.venv/bin/pytest tests/unit/test_demoad.py
See Add a New Model for the 5-minute walkthrough.
8. Check service health¶
exa status
curl http://localhost:18099/api/health
curl http://localhost:18001/health # Ray Serve (lists every loaded alias)
curl http://localhost:18002/health # Control plane (Phase 4)
exa modelzoo status # Phase 12 — freshness badges
exa modelzoo events # recent ModelZoo push events
ModelZoo freshness badges on the dashboard Models page will show CURRENT or UPDATED depending on whether any ModelZoo repository pushes have been recorded since the last retrain.
9. Start the notebook environment (Phase 9)¶
Opens JupyterHub at http://localhost:18888. Log in with your credentials. Each user gets an isolated JupyterLab container pre-configured to reach all stack services (MLflow, MinIO, Ray Serve, Prefect) by internal hostname.
VS Code: Jupyter extension → Specify Jupyter Server → http://localhost:18888/user/<username>/?token=<token>
See JupyterHub Guide for user management and full setup details.
Next steps¶
- Browse all UIs and APIs → Interfaces guide
- Add your own model → Add a new model
- Architecture overview → Architecture
- Trigger retrains from clients → Control Plane
- Notebook environment → JupyterHub Guide
- Add a non-sklearn framework → Add a new model
- Connect to an HPC cluster → Distributed training