MLGatee
How it works

How MLGatee
is built.

Three parts, one rule: your model file is never loaded on our own servers. Here is the whole path, from upload to answer.

Architecture

Three parts.

Your side, the control plane, and one runtime service per model.

1 · Your side

You

  • Web appapp.mlgatee.com
  • Python clientmlgatee.deploy()
  • Any HTTP clientcurl, Python, your app
2 · Control plane

MLGatee

  • Sign-inGitHub, Google, email link
  • Private storagemodel files, newest 3 versions
  • Byte inspectorreads bytes, never loads
  • Deploys and versionskeys stored as hashes
  • Monitoringcounts and timings only
3 · Model runtime

Your model’s service

  • One service per modelnothing shared
  • Python 3.12FastAPI server
  • Exact packagespinned from your file
  • /predict/predict_proba for classifiers
POST /predict with your key goes straight to your model’s own address. MLGatee isn’t in the request path.
Deploy flow

From upload
to live.

What happens between choosing a file and getting your endpoint.

Statusuploading
UploadingInspectingDeployingLive

Typically 1.5 to 4 minutes

  1. 1

    You start a deploy

    The web app or the Python client asks for a single-use upload link. Your plan’s limits are checked first: one model on the free plan, 50 MB per file.

  2. 2

    The file goes straight to private storage

    Your browser or notebook uploads the file directly to a private bucket. It is stored, never opened as a program.

  3. 3

    The bytes are inspected

    The inspector reads which libraries and versions the model needs and checks it is safe, without loading it. If something is wrong you get a plain message and nothing builds.

  4. 4

    Your model gets its own service

    A new service is created just for this model, with the exact packages pinned and only a hash of its key.

  5. 5

    Packages install, the model starts

    The service installs the packages, fetches your file through a short-lived signed link and starts a Python 3.12 server.

  6. 6

    You can watch it happen

    The status moves from uploading to inspecting to deploying, and every step appears in the Activity log.

  7. 7

    It’s live

    You get the endpoint address and a key starting mlg_live_, shown once. Typical time from upload: 1.5 to 4 minutes.

The byte inspector

Everything it needs
is in the bytes.

Pickles store the modules they need as plain text, and libraries like scikit-learn and XGBoost store their version. So the inspector can pin exact packages without ever loading the model.

Found in the fileInstallsVersion
sklearnscikit-learnYour file’s version (else 1.9.1)
xgboost (pickle, joblib)xgboost-cpuYour file’s exact version (2.1.1 or newer; else 3.4.1)
XGBoost .json / .ubjxgboost-cpu + scikit-learnYour file’s version or 3.4.1, whichever is newer
lightgbmlightgbm4.7.0
catboostcatboostYour file’s version or 1.2.10, whichever is newer
pandaspandas3.0.6
numpynumpy2.5.3, or numpy<2 for numpy 1 files
.onnxonnxruntime1.24.4
.tfliteai-edge-litert2.2.0
Hugging Face text (packed ONNX)onnxruntime + tokenizers1.24.4 · 0.23.2

Refused, with a plain message

Notebook classes

Models that need a class defined in your notebook. The message names the class and how to fix it.

Your own .py code

Classes or functions from your own files, which the server wouldn’t have.

Files that run commands

Anything that would call os.system, eval, subprocess and similar when loaded.

Some compression

bz2, xz and lz4 compressed joblib files. zlib and gzip work.

Zip bombs and damage

Anything unpacking past 200 MB, damaged or cut-off files.

Not yet supported

Native PyTorch, TensorFlow SavedModel and full Hugging Face Transformers. Convert them with the Python client.

All models run on Python 3.12. Every package pin is checked against a strict pattern before it reaches the build, so nothing else can be slipped into the build command.

Updating a model

New versions without
a new address.

Replacing a model keeps its address and key. Here is how.

1

Start

Upload a new file for a live model, or deploy the same name again from Python.

2

Check

The new file goes through the same inspection as a first deploy. A refused file doesn’t use up a version number.

3

Apply

The service rebuilds with the new file. The current version keeps answering; if the build fails, it stays live.

4

Swap

Once the new version is live it takes over, at the same address with the same key.

Rollback is a new version. MLGatee keeps the files of your three newest versions. Rolling back to one builds a new version from that file, so history only moves forward.

Calling a model

Rows in.
Predictions out.

Send a list of rows: numbers in training order, or objects with column names for table models.

Terminal
export MLGATEE_KEY="mlg_live_…"

curl -X POST https://mlg-iris-7f3a9c21.onrender.com/predict \
  -H "Authorization: Bearer $MLGATEE_KEY" \
  -H "Content-Type: application/json" \
  -d '{"input_data": [[5.1, 3.5, 1.4, 0.2]]}'
# → {"output":[0],"model":"iris"}
Python, not for Terminal
import os, requests

res = requests.post(
    "https://mlg-iris-7f3a9c21.onrender.com/predict",
    headers={"Authorization": f"Bearer {os.environ['MLGATEE_KEY']}"},
    json={"input_data": [[5.1, 3.5, 1.4, 0.2]]},
)
print(res.json())  # {'output': [0], 'model': 'iris'}

/predict

Returns {"output": [0], "model": "iris"}. “model” is your deployment’s name, even after a switch to another framework.

/predict_proba

Classifiers also return probabilities, like [[0.99, 0.005, 0.003]].

Text models

Hugging Face text models take plain text: ["great product!"].

Wrong key

A wrong or missing key gets {"error": "invalid or missing API key"}.

Monitoring

Counts, not contents.

Each service counts its own calls and reports about once a minute.

Recorded

  • Calls, per hour and per version
  • Answered, bad request, wrong key and server error counts
  • Answer times, as typical and slowest 5%
  • Rows predicted and the last call

Never recorded

  • Inputs you send
  • Outputs the model returns
  • Keys and request headers
Built with

The stack.

Four parts, each doing one job. The Python client is open source.

Web app

Next.js (React, TypeScript) on Cloudflare. Deploy, versions, rollback, monitoring, keys and access tokens.

Backend

Supabase: sign-in with GitHub, Google or an email link, Postgres with row-level security, a private storage bucket, and edge functions written in TypeScript.

Model runtime

Render, one web service per model. Python 3.12 with FastAPI. Free models use 512 MB and sleep after 15 idle minutes.

Python client

mlgatee 0.8.0, open source under Apache-2.0. Standard library only, Python 3.9 and newer.

Get started

Your model deserves
to see the real world.

Upload a file or deploy from your notebook. MLGatee handles the packages, the server and the key.

Free plan: one model, no card needed