Udgam-384M-Base

A 384M-parameter English base model trained from scratch by UdgamLabs. It is a decoder-only transformer compatible with the Gemma 4 architecture (Gemma4ForCausalLM), so it runs unmodified in 🤗 transformers, llama.cpp / Ollama / LM Studio (GGUF) and MLX. It is the foundation of the on-device tool-calling model Udgam-384M-Chat. For a 4,096-token context use Udgam-384M-Base-4K.

This is a base (pretrained) model: it continues text. It is not instruction-tuned and not safety-tuned, and it has no chat behaviour of its own. For chat and tool calling use Udgam-384M-Chat.

Model family: Udgam-384M-Base (2,048 context) · Udgam-384M-Base-4K · Udgam-384M-Chat · GGUF · MLX · MLX 4-bit · Technical report

Developer UdgamLabs
Parameters 384.4M total (317.3M in the 20 transformer layers; untied 32,768 × 1,024 input and output embeddings)
Architecture decoder-only transformer compatible with Gemma 4 (Gemma4ForCausalLM): 20 layers × 1,024, 16 query / 4 KV heads, sliding-window (512) and global attention in a 3 : 1 pattern, GeGLU MLP, logit soft-cap
Vocabulary 32,768 tokens, with the Gemma 4 chat-template special tokens
Context 2,048 tokens (max_position_embeddings 2048; GGUF context_length 2048)
Language English
Training pretrained from scratch on roughly 8B tokens of licensed public and synthetic text
Licence Apache-2.0

How to use

from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("UdgamLabs/Udgam-384M-Base")
model = AutoModelForCausalLM.from_pretrained("UdgamLabs/Udgam-384M-Base")
ids = tok("<bos>The capital of France is", return_tensors="pt", add_special_tokens=False).input_ids  # add <bos> yourself
print(tok.decode(model.generate(ids, max_new_tokens=30, do_sample=False)[0]))

As with official Gemma 4, tokenizer(text) does not add <bos>; prepend it (or use the chat template).

GGUF (in gguf/: f16 and Q8_0) with llama.cpp: llama-cli -m Udgam-384M-Base-Q8_0.gguf -c 2048 -p "The capital of France is" -n 40. Ollama (raw completion):

FROM ./Udgam-384M-Base-Q8_0.gguf
TEMPLATE """{{ .Prompt }}"""
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER repeat_penalty 1.15
PARAMETER num_ctx 2048

MLX: mlx_lm.generate --model UdgamLabs/Udgam-384M-Base --ignore-chat-template --prompt "<bos>The capital of France is" (mlx-lm reads the safetensors directly). Keep the literal <bos>: without it the model degenerates ("is is is…"); without --ignore-chat-template the prompt is wrapped in the chat template and the base model rambles.

Export parity was checked: transformers fp32, llama.cpp GGUF and MLX give the same greedy tokens, and Ollama's greedy output matches transformers.

Evaluation

Held-out bits per byte at 2,048 tokens (lower is better): 0.668 on mixed validation text, 0.865 on web text, 0.408 on tool-call / JSON / chat-style text.

Probes (greedy, raw prompts in the Gemma 4 format): a set_timer tool prompt gives a correct call <|tool_call>call:set_timer{label:<|"|>pasta<|"|>,minutes:15}<tool_call|>; a shell prompt gives find /var/log -type f -size +10M; "The chemical symbol of gold is" → " Au." Typical base-model weaknesses: repetition ("The capital of France is Paris, and the capital of Belgium is Brussels. The capital of France is …"), date-reasoning errors, buggy code. The pretraining text included structured and conversational material, so the base already knows the Gemma 4 tool-call format: prompted with the chat template, the 4K base scores BFCL v3 simple 66% / multiple 75% but irrelevance only 11% (it calls a tool almost always). We did not run knowledge benchmarks (HellaSwag, ARC, MMLU); at 384M parameters and ~8B tokens expect scores in line with other sub-500M models trained on similar budgets, far below larger models.

Limitations

  • Small model, small token budget: limited world knowledge, frequent factual errors, repetition, date-reasoning errors and buggy code.
  • English only. Context 2,048 tokens (use Udgam-384M-Base-4K for longer inputs).
  • Not instruction-tuned or safety-tuned: it will continue any text, including harmful text. Do not deploy it directly to end users.
  • Prompted as a chatbot without fine-tuning, it tends to ramble in a "thinking aloud" style ("Okay, let me…", "Wait, …"). Fine-tuning (as in Udgam-384M-Chat) removes this.
  • Parts of the training text were machine-generated; their styles and errors can carry over.

Training

Pretrained from scratch on roughly 8B tokens of licensed public and synthetic text, at a sequence length of 2,048. See the technical report for an overview.

Licence

Apache-2.0 (weights and tokenizer); see LICENSE. The chat template is the official Gemma 4 template (Apache-2.0).

Citation

@misc{udgamlabs2026udgam384m,
  title        = {Udgam-384M: a small on-device tool-calling model trained from scratch},
  author       = {{UdgamLabs}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/UdgamLabs/Udgam-384M-Chat}},
  note         = {Technical report: TECHNICAL_REPORT.md in the Udgam-384M-Chat repository}
}

Contact

UdgamLabs. Please use the Community (discussions) tab of the model page for questions, issues and feedback.

Downloads last month
7
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for UdgamLabs/Udgam-384M-Base

Quantizations
1 model