A Django REST service for credit-default risk scoring that treats the model as a deployed artifact rather than a pickle behind a route: algorithms are registered with a version and a status, predictions are logged, and two models can be A/B tested against live traffic and promoted on the result.
Built in 2022 for DiiMO at Elaniin. The classifier answers one question — will this applicant default? — and returns a probability plus a label.
Reference implementation, not a hardened deployment. It ran against SQLite with
DEBUG = True. See Caveats before copying anything from it.
Most "ML API" examples load one model at import and serve it forever. Replacing the model means a redeploy, comparing two models means guessing, and nobody can tell you what the API predicted last Tuesday.
This one puts a registry in front of the models instead:
Endpoint "default_risk_model"
├── MLAlgorithm "random forest" v1.0 ── status history ──► production
└── MLAlgorithm "extra trees" v1.0 ── status history ──► testing
│
└── every call writes an MLRequest row
(input, full response, label, feedback)
An algorithm's status is a record, not a column — MLAlgorithmStatus rows
accumulate and only the newest is active. So promotion and rollback are both
just writing a new status, and the history of what served production is
queryable after the fact.
POST /api/v1/default_risk_model/predict?status=production
Content-Type: application/json
{ "applicant_income": 5849, "loan_amount": 146, "credit_history": 1, ... }{ "probability": 0.83, "label": "Default", "status": "OK", "request_id": 55 }status selects which algorithm serves the call (production, testing,
staging, ab_testing); version pins a specific one. If more than one
algorithm matches and you are not A/B testing, the API returns 400 asking you
to specify a version rather than picking one silently — an ambiguous
production route is a bug, not a default.
The returned request_id is the handle for sending ground truth back later.
Start a test between two registered algorithms:
POST /api/v1/abtests # both algorithms move to status "ab_testing"
POST /api/v1/default_risk_model/predict?status=ab_testingTraffic splits 50/50 per request. Every prediction is stored as an MLRequest
with an empty feedback field; you fill that in as real outcomes arrive
(PATCH /api/v1/mlrequests/<id>), which is what makes the test scoreable at all
— a default takes months to observe, so accuracy cannot be computed at
prediction time.
POST /api/v1/stop_ab_test/<id>Counts, for each algorithm, the requests in the test window where response
matches feedback, promotes the more accurate one to production and demotes
the other to testing, and returns both accuracies.
| Method | Path | Does |
|---|---|---|
POST |
/api/v1/<endpoint>/predict |
Score one applicant |
GET |
/api/v1/endpoints |
Registered endpoints |
GET |
/api/v1/mlalgorithms |
Algorithms, versions, owners, source |
GET |
/api/v1/mlalgorithmstatuses |
Status history |
GET/PATCH |
/api/v1/mlrequests |
Logged predictions; PATCH to add feedback |
POST |
/api/v1/abtests |
Open an A/B test |
POST |
/api/v1/stop_ab_test/<id> |
Close it and promote the winner |
Trained in research/ (diimo.ipynb, AB_Testing_DiiMO-Default_Risk.ipynb) and
serialised with joblib:
random_forest.joblib— served as productionextra_trees.joblib— served as testingtrain_mode.joblib— column modes, used to fill missing fields at inference so training and serving impute the same wayregressor.joblib,encoders.joblib— supporting artifacts
Each model is wrapped in a class with the same four methods —
preprocessing / predict / postprocessing / compute_prediction — so the
registry can hold any of them without knowing what is inside. Adding a third
model is a new wrapper plus a registry.add_algorithm(...) line in wsgi.py.
cd backend/server
pip install -r ../requeriments.txt
python manage.py migrate
python manage.py runserverModel artifacts load from a path relative to the working directory, so start the
server from backend/server or the registry comes up empty (it catches the
failure and logs it rather than crashing the app).
With Docker — Django behind gunicorn, nginx in front:
docker-compose up --buildTests:
python manage.py test apps.ml.tests # model wrappers
python manage.py test apps.endpoints.tests # API surfaceHonest list, because this is a 2022 reference build and not a template to deploy:
SECRET_KEYis committed andDEBUG = True. Development settings. Move both to the environment before this faces anything real, and rotate the key.db.sqlite3is committed with the demo registry and ~54 logged predictions in it. Useful for reading, wrong for running.- The A/B split is per-request, not sticky. Fine here — scoring is a stateless one-shot call with no user session to keep consistent — but do not copy the pattern into anything where a user should see the same variant twice.
- Stopping a test with zero requests divides by zero. It assumes the test actually received traffic.
- Unlabelled requests count as wrong. Accuracy compares
responsetofeedback, andfeedbackdefaults to"", so anything you never labelled drags both numbers down. Compare the two algorithms to each other, not to an absolute bar. - Django 2.2 era (
django.conf.urls.url, Python 3.6 bytecode in the tree). It will not run unmodified on current Django.
The registry / status-history / request-log / A/B-promotion loop is the part that generalises. It is the difference between shipping a model and operating one, and it is the same shape whether the thing behind the endpoint is a random forest or an LLM — versioned, observable, swappable, and measurable after the fact against outcomes that arrive late.
Built by David Quintanilla — full stack and agentic AI.