Instructions to use PleIAs/Baguettotron with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PleIAs/Baguettotron with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PleIAs/Baguettotron") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("PleIAs/Baguettotron") model = AutoModelForCausalLM.from_pretrained("PleIAs/Baguettotron", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PleIAs/Baguettotron with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PleIAs/Baguettotron" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/Baguettotron", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PleIAs/Baguettotron
- SGLang
How to use PleIAs/Baguettotron with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PleIAs/Baguettotron" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/Baguettotron", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PleIAs/Baguettotron" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/Baguettotron", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PleIAs/Baguettotron with Docker Model Runner:
docker model run hf.co/PleIAs/Baguettotron
Incorrect eos token
Thanks for your work on this model and for sharing it openly!
The docs specify a prompt format like this:
<|im_start|>user
Who are you?<|im_end|>
<|im_start|>assistant
<think>
However, it looks like the tokenizer defines the eos token like this (note the lack of trailing pipe):
"65492": {
"content": "<|im_end>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
When using the provided prompt format, <|im_end|> is tokenized to multiple tokens, and generation uses multiple tokens to render <|im_end|>:
[
("<|", 45720),
("im", 317),
("_", 66),
("end", 416),
("|>", 31147)
]
This doesn't seem ideal, but I'm not sure how it could be best resolved.
Ok and yes not optimal. Won't really affect generation aside from costing for tokens since the model was trained with this sequence anyway.
We might retrain Baguettotron for the paper, so might be the occasion to fix it.
That makes sense. This may also cause issues with generation tools that expect to stop on a specific eos token rather than a stop string.
For what it's worth, Monad appears to include a similar choice, except Monad doesn't appear to include <|im_start|> or <|im_end|> as separate tokens in its vocabulary.
I wrote a script to help with this, see below.
⚠ Note that it was generated with me+gpt5+the readme/blog in this repo, would be great if someone else could check/correct this!
"""Utility to patch Baguettotron tokenizer w a proper chat template .jinja
Key improvements vs prior quick script:
* Avoid leading indentation/newlines inside rendered template
* Allow local path override
* Validate special tokens & show a sample rendering
* Optional disable thinking flag during demonstration
Run:
python scripts/fix_chat_template.py --source ./PleIAs_Baguettotron --out ./baguettotron_fixed
"""
from __future__ import annotations
import argparse
import json
from pathlib import Path
from typing import Any, Dict
from transformers import AutoTokenizer
# Compact, whitespace–controlled template matching README semantics.
# We deliberately avoid extra indentation to keep tokenization stable.
CHAT_TEMPLATE = (
"{%- for message in messages %}" # loop messages
"<|im_start|>{{ message['role'] }}\n" # role header
"{{ message['content'] }}" # content
"{%- if message['role'] == 'user' and message.get('sources') %}" # RAG sources
"{%- for source in message['sources'] %}\n<source_{{ loop.index }}>{{ source }}</source_{{ loop.index }}>{%- endfor %}"
"{%- endif %}" # end sources
"<|im_end|>\n" # end message
"{%- endfor %}" # end loop
"{%- if add_generation_prompt %}" # gen prompt
"<|im_start|>assistant\n"
"{%- if disable_thinking %}</think>\n{%- else %}<think>\n{%- endif %}"
"{%- endif %}" # end gen prompt
)
def ensure_special_tokens(tokenizer) -> None:
"""Guarantee bos/eos/pad tokens exist; don't clobber existing ids."""
# If pad is missing, prefer adding [PAD]; else fallback to eos.
if tokenizer.pad_token_id is None:
if "[PAD]" not in tokenizer.get_vocab():
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
else:
tokenizer.pad_token = "[PAD]"
# Avoid changing existing bos/eos if present.
if tokenizer.bos_token_id is None:
tokenizer.bos_token = "<|im_start|>"
if tokenizer.eos_token_id is None:
tokenizer.eos_token = "<|im_end|>"
def patch_config(config_path: Path, template: str) -> None:
data: Dict[str, Any] = json.loads(config_path.read_text())
data["chat_template"] = template
# Only set textual tokens if absent (no id assignment here)
data.setdefault("bos_token", "<|im_start|>")
data.setdefault("eos_token", "<|im_end|>")
data.setdefault("pad_token", "[PAD]")
config_path.write_text(json.dumps(data, indent=2, ensure_ascii=False))
def demo_render(tokenizer) -> str:
sample = [
{"role": "user", "content": "What is 3+5?"},
{"role": "user", "content": "Add a second question."},
]
return tokenizer.apply_chat_template(
sample, add_generation_prompt=True, tokenize=False, disable_thinking=False
)
def main():
ap = argparse.ArgumentParser(description="Fix Baguettotron chat template")
ap.add_argument(
"--source", default="./PleIAs_Baguettotron", help="Model dir or HF repo"
)
ap.add_argument("--out", default="./baguettotron_fixed", help="Output directory")
args = ap.parse_args()
tokenizer = AutoTokenizer.from_pretrained(args.source)
tokenizer.chat_template = CHAT_TEMPLATE
ensure_special_tokens(tokenizer)
out_dir = Path(args.out)
tokenizer.save_pretrained(str(out_dir))
patch_config(out_dir / "tokenizer_config.json", CHAT_TEMPLATE)
rendered = demo_render(tokenizer)
print("✓ Saved fixed tokenizer ->", out_dir)
print(f"✓ Template length: {len(CHAT_TEMPLATE)} chars")
print("--- Demo Render (truncated) ---")
print(rendered[:220] + ("..." if len(rendered) > 220 else ""))
if __name__ == "__main__":
main()
which gets me
{%- for message in messages %}
{{- '<|im_start|>' + message['role'] + '\n' }}
{{- message['content'] }}
{%- if message['role'] == 'user' and message.get('sources') %}
{%- for source in message['sources'] %}
{{- '\n\n<source_' + loop.index|string + '>' + source + '</source_' + loop.index|string + '>' }}
{%- endfor %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- if disable_thinking %}
{{- '</think>\n' }}
{%- else %}
{{- '<think>\n' }}
{%- endif %}
{%- endif %}
I can also load the corrected tokenizer from disk, made by ^ and push as a PR if you would like @pclanglais