AIWAZIRI · N-ATLAS DEVELOPER GATEWAY
Documentation
1. Introduction
N-ATLaS is a Nigerian-language large language model published as open weights on Hugging Face (NCAIR1/N-ATLaS). It is a Llama-3 8B model fine-tuned by Awarri Technologies with the Federal Ministry of Communications, Innovation and Digital Economy. The model card lists support for English, Hausa, Igbo and Yoruba.
Using open weights directly means provisioning a GPU, loading about 16 GB of weights, running an inference server and writing integration code. The AIWAZIRI N-ATLAS Developer Gateway does that work once and exposes N-ATLaS through a documented HTTP API, a Python SDK, a browser playground and a multilingual comparison tool.
AIWAZIRI builds and operates the gateway. AIWAZIRI does not own or train N-ATLaS.
2. Quick start
1 — Check the service
curl https://natlas.aiwaziri.com/api/v1/healthinference is available when the GPU worker is serving N-ATLaS, and starting during a cold start. Cold starts take a few minutes because the GPU scales to zero when idle.
2 — Send a chat request
curl -X POST https://natlas.aiwaziri.com/api/v1/chat \
-H "Content-Type: application/json" \
-d '{
"model": "N-ATLaS",
"language": "hausa",
"messages": [{ "role": "user", "content": "Menene basirar wucin gadi?" }]
}'3 — Or use Python
pip install "git+https://github.com/AIWAZIRI-LIMITED/aiwaziri-natlas#subdirectory=packages/python-sdk"
from aiwaziri_natlas import NAtlas
client = NAtlas(base_url="https://natlas.aiwaziri.com")
print(client.chat(message="Menene basirar wucin gadi?", language="hausa").text)3. N-ATLAS integration
The gateway runs the official N-ATLaS weights. There is no other model behind it, and no responses are cached or pre-written.
| Model | NCAIR1/N-ATLaS, pinned to revision e294476 |
|---|---|
| Architecture | LlamaForCausalLM (Llama-3 8B), 32 layers, 128,256-token vocabulary |
| Precision | bfloat16, the published dtype. No quantisation. |
| Inference server | vLLM (OpenAI-compatible) using the model's own Llama-3 chat template |
| Hardware | 1× NVIDIA L40S (48 GB) on Modal. Minimum is a 24 GB GPU for bf16 at 8k context. |
| Context | 8,192 tokens (the model card recommends ≤ 8,092 for best performance) |
| Sampling defaults | temperature 0.6, top_p 0.9 (from generation_config.json); repetition_penalty 1.12 (from the model card example) |
| Weights access | Downloaded at deploy time with a Hugging Face token from an account that accepted the N-ATLaS terms |
Integration method: self-hosted model weights. The gateway does not call an official N-ATLaS API, because none was available to the project. The gateway only needs an OpenAI-compatible endpoint (NATLAS_UPSTREAM_URL), so an official N-ATLaS API can be connected later without changing the public API.
4. REST API
Base URL: https://natlas.aiwaziri.com/api/v1. Requests and responses are JSON (UTF-8). Every response has an X-Request-ID header. You can send your ownX-Request-ID to correlate logs.
GET/api/v1/health
Probes the GPU worker and reports whether N-ATLaS is actually being served. Results are cached for 20 s. The gateway always answers with HTTP 200 and puts the state in inference.
{
"status": "ok", // "ok" | "degraded"
"gateway": "ok",
"model": "N-ATLaS",
"model_id": "NCAIR1/N-ATLaS",
"inference": "available", // "available" | "starting" | "unavailable" | "unconfigured"
"served_models": ["N-ATLaS"],
"probe_ms": 41
}GET/api/v1/models
Lists the models available through the gateway. Currently that is only N-ATLaS. The response includes serving details, license and attribution.
curl https://natlas.aiwaziri.com/api/v1/modelsGET/api/v1/languages
Lists the supported language identifiers: english, hausa, yoruba, igbo.
POST/api/v1/chat
Multi-turn chat completion with N-ATLaS.
| Field | Type | Notes |
|---|---|---|
messages | array, required | 1–32 items of {role, content}. Roles are system, user or assistant. The last message must be user. Content is 1–6,000 characters. |
language | string | english (default), hausa, yoruba or igbo. Adds a system instruction to respond in that language, unless you send your own system message. |
model | string | N-ATLaS (default) |
max_tokens | int | 1–1024, default 512 |
temperature | float | 0–2, default 0.6 |
top_p | float | (0, 1], default 0.9 |
stream | bool | When true, the response is sent as server-sent events (see below) |
{
"request_id": "req_3f9c0d1e2a4b5c6d7e8f",
"model": "N-ATLaS",
"model_id": "NCAIR1/N-ATLaS",
"language": "hausa",
"response": "…generated by N-ATLaS…",
"finish_reason": "stop",
"latency_ms": 2140,
"usage": { "prompt_tokens": 61, "completion_tokens": 118, "total_tokens": 179 }
}latency_ms is measured by the gateway around the inference call. usage holds the token counts vLLM reports, and is null if vLLM did not report them.
Streaming (stream: true)
The response is text/event-stream with these events:
event: start
data: {"request_id":"req_…","model":"N-ATLaS","model_id":"NCAIR1/N-ATLaS","language":"igbo"}
event: delta
data: {"text":"Ndewo"}
event: done
data: {"request_id":"req_…","finish_reason":"stop","latency_ms":3012,"time_to_first_token_ms":410,"usage":{…}}If inference fails after the stream has started, you receive event: error with {code, message, request_id}.
curl -N -X POST https://natlas.aiwaziri.com/api/v1/chat -H "Content-Type: application/json" \
-d '{"language":"yoruba","stream":true,"messages":[{"role":"user","content":"Kini oye atọwọda?"}]}'POST/api/v1/generate
Single-prompt shortcut. Takes prompt (string, required) in place of messages. Every other field and the response shape are the same as/chat. The prompt is sent as one user turn through the model's chat template.
curl -X POST https://natlas.aiwaziri.com/api/v1/generate -H "Content-Type: application/json" \
-d '{"prompt":"Gịnị bụ ọgụgụ isi arụrụ arụ?","language":"igbo","max_tokens":200}'The gateway's own OpenAPI schema is served at /docs and /openapi.json on the gateway origin.
5. Python SDK
The aiwaziri-natlas package (Python ≥ 3.9, depends only on httpx) is not yet on PyPI. Install it from the repository:
pip install "git+https://github.com/AIWAZIRI-LIMITED/aiwaziri-natlas#subdirectory=packages/python-sdk"from aiwaziri_natlas import NAtlas, Message, NAtlasError
client = NAtlas(base_url="https://natlas.aiwaziri.com") # or set NATLAS_BASE_URL
r = client.chat(message="Menene basirar wucin gadi?", language="hausa")
r.text, r.latency_ms, r.request_id, r.usage, r.finish_reason
# streaming
for ev in client.stream(message="Kini oye atọwọda?", language="yoruba"):
if ev.type == "delta":
print(ev.text, end="", flush=True)
# multi-turn
client.chat(messages=[Message("user", "Ndewo!")], language="igbo")
# single prompt, health, models
client.generate("Explain photosynthesis simply.")
client.health(); client.models()The package installs a CLI:
natlas "Menene basirar wucin gadi?" --language hausa
natlas --health6. Playground & Language Lab
- Playground: multi-turn streaming chat with N-ATLaS. Shows latency, time to first token, token counts, finish reason and request ID for each reply. A Stop button cancels a request. A ready-to-run curl command mirrors the current conversation.
- Language Lab: runs one prompt in English, Hausa, Yorùbá and Igbo at the same time and shows the outputs side by side for comparison.
7. Supported languages
| ID | Language | Model card human-eval average* |
|---|---|---|
english | English | 4.21 / 5 (n=1,662) |
hausa | Hausa | 3.98 / 5 (n=140) |
igbo | Igbo | 3.87 / 5 (n=296) |
yoruba | Yorùbá | 2.69 / 5 (n=542) |
*These figures are quoted from the official N-ATLaS model card. They are not AIWAZIRI measurements. The model card itself notes that Yoruba has room for improvement. AIWAZIRI has not benchmarked the model.
8. Authentication
The public beta does not need an API key, and rate limits protect it instead (20 requests per minute per IP per gateway instance). Operators can require keys by setting GATEWAY_API_KEYS. Clients then send X-API-Key: <key> or Authorization: Bearer <key>. Health, models and languages stay open. The gateway's credential for the GPU worker is stored only as a Modal secret and is never sent to clients.
9. Error handling
{ "error": { "code": "invalid_request", "message": "Request validation failed.", "request_id": "req_…", "details": [ … ] } }| HTTP | code | Meaning / action |
|---|---|---|
| 422 | invalid_request | Body failed validation. details lists each field and what was wrong with it. |
| 413 | payload_too_large | Body is larger than 64 KB |
| 401 | unauthorized | API key missing or invalid (only when keys are enabled) |
| 429 | rate_limited | Wait and retry |
| 503 | inference_starting | GPU worker is starting or busy. Retry with backoff. |
| 504 | inference_timeout | Inference took too long, usually during a cold start. Retry. |
| 502 | inference_unreachable / inference_auth / inference_error | Backend problem. Report the request_id. |
| 400 | inference_rejected | vLLM rejected the request, usually because the prompt is longer than the context window |
Error messages are sanitised. Details from the upstream service, including credentials, are never returned to clients.
10. Examples
JavaScript (fetch)
const res = await fetch("https://natlas.aiwaziri.com/api/v1/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ language: "yoruba", messages: [{ role: "user", content: "Kini oye atọwọda?" }] }),
});
const { response, latency_ms } = await res.json();Python with retry on cold start
import time
from aiwaziri_natlas import NAtlas, NAtlasError
client = NAtlas()
for attempt in range(4):
try:
print(client.chat(message="Ka fassara 'good morning' zuwa Hausa.", language="hausa").text)
break
except NAtlasError as e:
if not e.retryable: raise
time.sleep(30 * (attempt + 1))The repository has more examples in packages/python-sdk/examples/: quickstart across all four languages, streaming, and multi-turn Igbo.
11. Deployment
The full guide is docs/deployment.md in the repository. In summary:
# 1. Inference + gateway on Modal (needs an HF token whose account accepted the N-ATLaS terms)
modal secret create huggingface HF_TOKEN=hf_...
modal secret create natlas-upstream NATLAS_UPSTREAM_KEY=$(openssl rand -hex 32)
modal deploy infra/modal/natlas_app.py
# 2. Web app on Cloudflare Pages
cd apps/web && npm ci && npm run build
npx wrangler pages deploy out --project-name aiwaziri-natlas
# set Pages env var GATEWAY_ORIGIN=https://<workspace>--natlas-gateway.modal.runTo self-host on your own GPU, run vllm serve NCAIR1/N-ATLaS --served-model-name N-ATLaS and point NATLAS_UPSTREAM_URL at it. Then start the gateway with uvicorn natlas_gateway.main:app.
12. Architecture
- Web (Cloudflare Pages): statically exported Next.js app. A Pages Function proxies
/api/*to the gateway on the same origin and streams SSE through unbuffered. - Gateway (FastAPI on Modal, CPU): validates requests, applies language prompts, enforces size limits and rate limits, optionally checks API keys, measures latency, handles streaming and sanitises errors.
- Inference (vLLM on Modal, GPU): serves N-ATLaS with continuous batching, behind a bearer token known only to the gateway. Weights are cached on a Modal Volume. The worker scales to zero after 15 minutes idle.
13. Limitations
- Cold starts: the first request after an idle period waits for a GPU and for the weights to load, which takes a few minutes. Later requests are fast until the worker scales down again.
- Language steering is prompt-based. The model may sometimes answer in English or mix languages, and the model card notes limited handling of code-switching.
- Quality varies by language, as the model card's own human evaluation shows. Do not use outputs for high-stakes decisions without human review.
- Context is limited to 8,192 tokens. Output is capped at 1,024 tokens per request.
- Rate limits are kept in memory for each gateway instance. There are no user accounts and no usage dashboard.
- The Python SDK is not yet on PyPI. There is no TypeScript SDK yet; use
fetchas shown above. - The N-ATLaS license caps use at 1,000 active end-users per 30 days. Commercial use needs a separate license from Awarri Technologies.
14. Attribution & license
N-ATLaS is attributed to: “Awarri Technologies and the Federal Ministry of Communications, Innovation and Digital Economy.” The model card also acknowledges NITDA, the National Center for Artificial Intelligence and Robotics (NCAIR), and data contributors from Nigeria's six geopolitical zones.
N-ATLaS is used under the N-ATLaS Terms of Use v1.0 (Open-Source Research and Innovation License). The gateway does not modify, fine-tune or redistribute the weights; it downloads them from Hugging Face at deploy time. AIWAZIRI's gateway, SDK and web code are released under Apache-2.0. See ATTRIBUTION.md in the repository.
15. Responsible use
The N-ATLaS Terms of Use prohibit, among other things:
- surveillance or unlawful monitoring
- discriminatory profiling
- disinformation, impersonation or synthetic fraud
- military or weaponised use
- exploitative or unlawful applications
Using the gateway means you agree to those terms. Model outputs can be wrong or biased; the model card reports bias concerns. Tell users when they are interacting with AI, and review outputs before acting on them. The gateway does not store prompts or responses. Operational logs record request IDs, language and latency only.
