Docs / Quickstart
Quickstart
From a trained model to a live, key-protected endpoint, then everything after: versions, rollback, keys and monitoring.
Get started
Sign in at app.mlgatee.com with GitHub, Google or an email link. Signing in with the same email address in different ways gives you one account. The free plan runs one model; there’s no card to enter.
Email links are single use and open in the browser that asked for them. If one says it expired, ask for a new one from the same browser.
Deploy from the web
Open Deploy a model (or press N), drop your model file and give it a name. Names use 3 to 40 characters: lowercase letters, numbers and dashes, starting with a letter. The name becomes part of your endpoint’s address.
Accepted files, up to 50 MB: .pkl and .joblib (scikit-learn, XGBoost, LightGBM, CatBoost), XGBoost’s own .json and .ubj, .onnx and .tflite.
- Upload. The file goes straight to private storage through a single-use link.
- Check. The file’s bytes are read, without loading it, to find the framework and the exact packages it needs. If something is wrong you get a plain message and nothing builds.
- Deploy. Your model gets its own service, with those packages pinned, on Python 3.12.
- Live. You get the endpoint address and a key starting
mlg_live_.
The key is shown once. Copy it into a password manager before you leave the page; MLGatee keeps only a hash and its last four characters. A deploy typically takes 1.5 to 4 minutes (CatBoost a little longer), and every step appears in the model’s Activity tab.
Deploy from Python
The Python client (version 0.8.0) uses only the standard library and works on Python 3.9 and newer.
pip install git+https://github.com/Reetcodeio/mlgatee-pythonIn the app, open Settings → Access tokens and create a token. Tokens start with mlg_pat_, last 90 days, and you can have 5 at a time. Then, in your notebook:
import mlgatee
mlgatee.login() # paste your access token once
mlgatee.deploy(model, name="iris") # prints progress, then the key oncelogin() asks for the token in a hidden prompt and saves it to ~/.mlgatee/token, readable only by you. You can set MLGATEE_TOKEN instead.
Before anything is uploaded, deploy() saves the model with joblib in your own Python and checks its size, the name, and classes defined in your notebook. Then it uses the same upload, check and build as the web app, prints progress, and prints the endpoint key once. The key is kept only in memory for the session, never on disk.
Call your model
Send a list of rows as input_data: numbers in training order, or objects with column names for models trained on a table, like [{"age": 30, "city": "Pune"}].
export MLGATEE_KEY="mlg_live_…"
curl -X POST https://mlg-iris-7f3a9c21.onrender.com/predict \
-H "Authorization: Bearer $MLGATEE_KEY" \
-H "Content-Type: application/json" \
-d '{"input_data": [[5.1, 3.5, 1.4, 0.2]]}'
# → {"output":[0],"model":"iris"}import os, requests
res = requests.post(
"https://mlg-iris-7f3a9c21.onrender.com/predict",
headers={"Authorization": f"Bearer {os.environ['MLGATEE_KEY']}"},
json={"input_data": [[5.1, 3.5, 1.4, 0.2]]},
)
print(res.json()) # {'output': [0], 'model': 'iris'}/predictreturns labels or values:{"output": [0], "model": "iris"}."model"is your model’s name, and stays the same across versions.- Classifiers also answer
/predict_proba:{"output": [[0.99, 0.005, 0.003]], ...}. - Hugging Face text models take text:
["great product!"]or{"text": "great product!"}. - A wrong or missing key gets
401 {"error": "invalid or missing API key"}; a body the model can’t use gets a 400.
Each model’s Usage tab has ready-to-run commands for Mac and Linux (curl), Windows PowerShell, Windows CMD, Python and JavaScript, with your address filled in. Calls go straight to your model’s own service; MLGatee isn’t in the request path.
New versions and rollback
Upload a new file on the model’s Versions tab, or deploy the same name again from Python. The new file is checked like a first deploy, then built while the current version keeps answering. Once it’s live it takes over, at the same address with the same key. If the build fails, the current version stays live. A refused file doesn’t use up a version number.
The files of your three newest versions are kept. Rolling back to one builds a new version from that file, so going from v3 back to v1 creates v4. History only moves forward.
mlgatee.versions("iris") # which versions you can roll back to
mlgatee.rollback("iris", 2) # builds a new version from v2's fileKeys, rebuild and delete
- New key (model → Settings): the model rebuilds with a new key, shown once. The old key keeps working until the rebuild is live, so nothing breaks in between.
- Rebuild: builds the model again from its file, with a new endpoint key (shown once). Use it after a failed build.
- Delete model: removes its service, its files and its record. The address stops answering.
Monitoring
Each model’s service counts its own calls and reports about once a minute. The Monitoring tab shows, for the last 24 hours or 7 days:
- Calls and rows predicted;
- Errors: bad requests and server errors, as a share of calls with a valid key;
- Answer time: typical and slowest 5%;
- Last called, an hourly chart with errors in red, and a By version table: the “should I roll back?” view.
Calls with a wrong key, and wake-ups after a nap, are noted separately. Inputs, outputs, keys and headers are never recorded.
mlgatee usage iris --window 7dKeras, PyTorch, Hugging Face
These are converted in your own Python, so they fit the standard 512 MB machine. With an example input, the client runs the converted model and checks its answers against the original before uploading (within 0.001, or 0.05 when quantized).
# Keras → TensorFlow Lite
mlgatee.deploy(keras_model, name="digits")
# PyTorch → ONNX (opset 18)
mlgatee.deploy(net, name="digits", example_input=X[:5], task="classifier")
# Hugging Face text model → one 8-bit ONNX file with its tokenizer
mlgatee.deploy_hf("philschmid/tiny-bert-sst2-distilled", name="reviews")- Keras becomes TensorFlow Lite with built-in ops; add
quantize=Truefor a smaller file. - PyTorch is exported to ONNX with a dynamic batch size.
task="classifier"adds labels and probabilities, so/predict_probaworks. - Hugging Face text models for classification or embeddings are packed with their tokenizer into one 8-bit ONNX file (
max_length128 by default).
Install the matching extra first: mlgatee[tensorflow], mlgatee[torch] or mlgatee[huggingface].
Plans and limits
| Free | Paid plans (when payments open) | |
|---|---|---|
| Models | 1 | Starter: 1 always on. Pro: 3 |
| File size | 50 MB | 50 MB |
| Builds | 5 per 24 hours, 20 per 30 days | Pro: 15 per 24 hours, 100 per 30 days |
| Machine | 512 MB, shared CPU | Always on; a Large 2 GB machine as an add-on |
| Idle | Sleeps after 15 quiet minutes; the next call wakes it in about a minute | Pro pauses after 48 idle hours; resume keeps the address and key |
A deploy, a new version, a rollback, a New key and a Rebuild each count as one build. Your builds today are shown at the bottom of the app’s sidebar. See pricing.
If a model is refused
The message says why in plain words, and nothing builds.
- A class from your notebook or your own
.pyfile. The message names it. Move it into an installed package, or rebuild the model from standard parts. - A file that would run commands when loaded (
os.system,eval,subprocessand similar). Refused for safety. - bz2, xz or lz4 compressed joblib files. Save with zlib or gzip compression instead.
- Files that unpack past 200 MB, damaged or cut-off files, and JSON that isn’t an XGBoost model.
- ONNX with custom operators or external data; TensorFlow Lite that needs full TensorFlow ops.
- Native PyTorch, TensorFlow SavedModel and full Hugging Face Transformers. Convert them with the Python client (above).
See the inspector rules for the packages and versions it installs.
Python and CLI reference
| Python · command line | Does |
|---|---|
login()mlgatee login | Paste an access token once; it’s saved to ~/.mlgatee/token |
deploy(model, name, plan=)mlgatee deploy FILE --name --plan | A new name is a new model; an existing name is its next version |
deploy_hf(model_id, name, task=) | A Hugging Face text model, packed into one file |
predict(name, rows, key=) | Call a model from Python |
models()mlgatee models | Your models |
versions(name)mlgatee versions NAME | Versions you can roll back to |
rollback(name, version)mlgatee rollback NAME V | Build a new version from an earlier file |
usage(name, window)mlgatee usage NAME --window 7d | Calls, errors and answer times |
plan()mlgatee plan | Your plan, limits and builds used |
resume(name)mlgatee resume NAME | Restart a paused Pro model: same address and key, no build |
delete(name)mlgatee delete NAME | Remove the model’s service, files and record |
logout()mlgatee logout | Forget the saved token |
Self-hosting? Point the client at your own project with MLGATEE_URL and MLGATEE_PUBLISHABLE_KEY.
Troubleshooting
| You see | What to do |
|---|---|
| The first call takes about a minute | A free model sleeps after 15 quiet minutes, and the next call wakes it. Expected. Paid plans keep models awake. |
invalid or missing API key (401) | The key is wrong, missing, or was replaced by a New key. Use the current key, or make a new one (shown once). |
| A 400 error | The body isn’t what the model expects. Send {"input_data": [...]} with rows in training order, or column names for table models. |
build_limit (429) | You’ve used 5 builds in 24 hours or 20 in 30 days. Wait for the window to pass; the message says when. |
storage_full | Uploads are paused while storage is nearly full. Models that are live keep answering. Try again later. |
| The build failed | Read the model’s Activity log. Most often a library couldn’t be installed, or the model didn’t fit in 512 MB. During a new version, the current one keeps answering. |
| Monitoring says “No reports yet” | The model was built before monitoring existed. A New key, a Rebuild or a new version turns it on. |
| Email link: “expired or opened in a different browser” | Open the link in the same browser you asked for it from, or ask for a new one. |
Still stuck? Send us a message.