Regnant
Research/R.01 · KW5-Lite

R.01The models, released when they are real

KW5-Lite.

A Kiswahili foundation model trained from scratch. Published so it can be checked, not because it is finished.

KW5-Lite is a 109.5M parameter decoder-only transformer trained from scratch on 2.1B Swahili tokens. It represents one of the first open-source foundation models built specifically for Kiswahili, with both base and instruction-tuned variants available. Trained on a single T4 GPU over 120 hours, it holds conversations in standard Kiswahili on modest hardware.

Trained from scratch

Unlike adapted or fine-tuned models, KW5-Lite was trained from the ground up on Swahili text. The base model learns Swahili at the weights, not through translation. The instruction-tuned variant adds conversational capabilities through efficient LoRA fine-tuning on 22.5K Swahili instruction-response pairs.

Base: 109.5M parameters · Instruct: base + LoRA adapter · Apache 2.0 · jump to the evaluation

Model card, as it stands

Class

Decoder-only transformer, trained from scratch

Parameters

109.5M trainable / 134M total with embeddings

Training Data

2.1B Swahili tokens from FineWeb-2 corpus

Architecture

12-layer transformer, RMSNorm, RoPE, SwiGLU

Variants

Base model + Instruction-tuned (LoRA fine-tuned)

License

Apache 2.0, open weights

Intended use

Research and further fine-tuning. Not for consequential work

Programme

R.01 LUGHA, sovereign language models

RUN 2ff422c0be23enterprise-swahili-kw5-lite-v2SEED 20260904

It compresses
Swahili better than
anything it was
measured against.

And it does not beat those models on the reasoning tasks. Both halves of that are the result. Nine models were run on one seeded pass over three Swahili benchmarks; the artifacts are linked at the foot of this section.

Bits per byte, held-out Swahili

1.201

Lowest in the field, from the smallest model in it, at 229MB of peak GPU memory. XGLM-564M needs 1.1GB to reach 1.336.

Parameters
109.5M
Peak GPU memory
229 MB
Throughput
74.5 tok/s

Leaderboard: PMI-normalized accuracy

ModelParamsbpb ↓BelebeleAfriXNLIMasakhaNEWStok/sPeak MBComposite
Random chance25.0%33.3%14.3%
KW5-Lite (base)RegnantSwahili specialist109.5M1.20126.1%32.5%23.1%74.52290.358
KW5-Lite (instruct)RegnantSwahili specialist109.5M1.30225.0%33.5%29.0%74.32250.564
swa-gpt2-100mbSwahili specialist124.8M1.28526.2%36.2%28.1%122.32570.733
XGLM-564MMultilingual compact564.5M1.33629.2%34.3%63.4%49.411180.905
BLOOM-560mMultilingual compact559.2M2.10528.1%33.8%41.2%49.311010.528
Qwen2.5-0.5BMultilingual compact494M2.96325.2%32.7%37.4%32.69660.449
Pythia-410mEnglish-only control405.3M2.78828.1%30.3%56.27980.648
GPT-2-124mEnglish-only control124.4M3.45527.7%21.6%105.62680.526
Gemma-3-270mNon-complete
  • KW5-Lite (instruct)Reported separately. The leaderboard is a base-model evaluation, not instruction following.
  • Pythia-410mAfriXNLI excluded: 99% of predictions landed on a single label.
  • GPT-2-124mAfriXNLI excluded: 84% of predictions landed on a single label.
  • Gemma-3-270mGated repository. Left in the table as non-complete rather than replaced or zeroed.
  • MethodAny model/task pair predicting one option for more than 80% of items is flagged as label-collapsed and excluded from the composite, because an accuracy figure produced by guessing the same answer every time is not an accuracy figure.

Reading the table

01

Two of the three tasks separate nobody

On Belebele every model in the field sits within a few points of the 25% chance line, and on AfriXNLI within a few points of 33%. Paired bootstrap intervals for KW5-Lite against every baseline on both tasks cross zero. No model here can read a Swahili passage and answer a question about it.

02

On the third, it is genuinely behind

MasakhaNEWS is where the intervals stop crossing zero. XGLM-564M reaches 63.4% and KW5-Lite base 23.1%, a gap of forty points that the bootstrap confirms. Five times the parameters and a hundred languages of pretraining buy real topic classification.

03

The compression result is the real one

Bits per byte measures how well the weights model Swahili text itself, and it is the one number comparable across tokenizers. KW5-Lite base is first at 1.201, ahead of XGLM at 1.336 and every English-pretrained model by a factor of two, with a fifth of XGLM's memory.

Figures, as the run produced them

FIG 01Accuracy by task, against chance
Grouped bar chart of PMI-normalized accuracy for every model on Belebele, AfriXNLI and MasakhaNEWS, with the random-chance line marked on each panel
FIG 02Paired bootstrap against every baseline
Forest plot of KW5-Lite base accuracy minus each baseline, with 95% bootstrap confidence intervals; Belebele and AfriXNLI intervals all cross zero, MasakhaNEWS intervals do not
FIG 03Bits per byte on held-out Swahili
Bar chart of bits per byte on held-out Swahili text; KW5-Lite base is lowest in the field
FIG 04Accuracy against parameter count
Accuracy plotted against parameter count, grouping models into size classes
FIG 05Throughput against peak memory
Generation throughput in tokens per second against peak GPU memory for each model
FIG 06Composite score
Composite score across all normalized metrics, ranked by model

01Where it holds up

A Kiswahili foundation model.

Use the instruction-tuned variant for conversational applications: casual question and answer, language practice, general knowledge that is well attested. The base model is for research and further fine-tuning.

  • 01Standard-register Kiswahili grammar, handled correctly across the evaluation set (instruction-tuned model)
  • 02Fluent text generation with proper Swahili syntax and subject-verb agreement (base model)
  • 03Familiar cultural and factual ground, where the answer is well attested in training data
  • 04Short exchanges and conversations, handling context across multiple turns (instruction-tuned model)
  • 05Runs efficiently on modest hardware (4GB GPU VRAM minimum, CPU inference possible)

02Where it fails

Four ways
it breaks.

These are on the published cards. A 109.5M parameter model trained on 2.1B tokens has limits, and printing them is cheaper for everyone than letting someone discover them in a production deployment.

01

Base model needs instruction tuning

The base model is trained for language modeling only. It doesn't follow instructions or engage in conversation without fine-tuning. Use the instruction-tuned variant for conversational applications.

02

Limited to standard Kiswahili

Trained primarily on standard Tanzanian Kiswahili. Casual phrasing, Sheng-influenced input, and other Swahili dialects make it noticeably less reliable.

03

Factual knowledge is limited

The 109.5M parameter size limits factual knowledge capacity. It will confidently generate plausible-sounding but factually incorrect information for topics outside its training data, especially current events and specialized domains.

04

Context window constraints

2048 token context window (roughly 1000 tokens effective in conversation). Long conversations or documents exceed this limit, causing the model to lose earlier context.

Do not use it for

Medical, legal, or financial advice. Any decision a person or an institution has to stand behind. Any workflow where a confident invented answer would be acted on.

Nothing in the Regnant portfolio depends on this model for consequential work. Where a system needs a decision it can defend, it is built to show its evidence, not to be trusted on fluency.

03What it is a step toward

Language equity
is a model problem,
not a translation one.

A translation layer over an English model inherits every place that model has nothing to say. KW5-Lite holds Kiswahili at the weights— trained from scratch on 2.1B Swahili tokens. It is small and limited, but it proves the approach. The point of publishing it is that the next one has something to beat.

Get it

Base model on Hugging Face ↗Instruction-tuned model on Hugging Face ↗

Base model: Standard PyTorch format, HuggingFace compatible. Instruct model: LoRA fine-tuned on 22.5K Swahili instruction pairs. Both models use modern architecture with RMSNorm, RoPE, and SwiGLU.

Related

Milkshake, the education ecosystem →The LUGHA programme →