Instructions to use RthItalia/Rth-lm-25b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use RthItalia/Rth-lm-25b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: llama cli -hf RthItalia/Rth-lm-25b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: llama cli -hf RthItalia/Rth-lm-25b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: ./llama-cli -hf RthItalia/Rth-lm-25b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: ./build/bin/llama-cli -hf RthItalia/Rth-lm-25b
Use Docker
docker model run hf.co/RthItalia/Rth-lm-25b
- LM Studio
- Jan
- Ollama
How to use RthItalia/Rth-lm-25b with Ollama:
ollama run hf.co/RthItalia/Rth-lm-25b
- Unsloth Desktop
- Docker Model Runner
How to use RthItalia/Rth-lm-25b with Docker Model Runner:
docker model run hf.co/RthItalia/Rth-lm-25b
- Lemonade
How to use RthItalia/Rth-lm-25b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull RthItalia/Rth-lm-25b
Run and chat with the model
lemonade run user.Rth-lm-25b-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
Update model card with FRO-LM Small v0
Browse files
reports/HF_RTH_LM_SWARMLM_V2_CARD.md
CHANGED
|
@@ -211,6 +211,55 @@ Current observed limitations:
|
|
| 211 |
|
| 212 |
These failures motivate targeted `orchestrator_v3` work, especially for code classification and FRO/text distinction.
|
| 213 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 214 |
## What This Release Supports
|
| 215 |
|
| 216 |
This release supports:
|
|
@@ -249,6 +298,10 @@ Recommended short description:
|
|
| 249 |
|
| 250 |
> RTH-LM / SwarmLM v2 is a modular Genome/Soul research system. A shared frozen Genome supports multiple rank-512 specialist Souls, while `orchestrator_v2` routes controlled requests to specialist executors. In a controlled cascade evaluation, SwarmLM v2 reached 87.5% route accuracy and 75% end-to-end cascade success.
|
| 251 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 252 |
## License
|
| 253 |
|
| 254 |
Model artifacts are released for research and non-commercial use under the project license unless a separate commercial license is granted by RTH Italia.
|
|
|
|
| 211 |
|
| 212 |
These failures motivate targeted `orchestrator_v3` work, especially for code classification and FRO/text distinction.
|
| 213 |
|
| 214 |
+
## Orchestrator v3b Routing Update
|
| 215 |
+
|
| 216 |
+
`orchestrator_v3b` is a targeted routing update trained after the v2 cascade evaluation. It keeps the same frozen Genome and the same v2 specialist Souls, but replaces the central routing checkpoint:
|
| 217 |
+
|
| 218 |
+
```text
|
| 219 |
+
orchestrator_v3b -> selected v2 specialist Soul -> output
|
| 220 |
+
```
|
| 221 |
+
|
| 222 |
+
Artifact path:
|
| 223 |
+
|
| 224 |
+
```text
|
| 225 |
+
souls/orchestrator_v3b/ORCHESTRATOR_V3B.pt
|
| 226 |
+
```
|
| 227 |
+
|
| 228 |
+
Controlled cascade result:
|
| 229 |
+
|
| 230 |
+
```text
|
| 231 |
+
Tasks: 8
|
| 232 |
+
Route accuracy: 1.000
|
| 233 |
+
Specialist marker score average: 0.875
|
| 234 |
+
Cascade success rate: 0.875
|
| 235 |
+
Average cascade latency: 48.36s
|
| 236 |
+
Average route tokens/sec: 17.72
|
| 237 |
+
Average specialist tokens/sec: 15.72
|
| 238 |
+
Peak route VRAM: 18.60 GB
|
| 239 |
+
Peak specialist VRAM: 18.62 GB
|
| 240 |
+
```
|
| 241 |
+
|
| 242 |
+
Task-level comparison:
|
| 243 |
+
|
| 244 |
+
| Task | Expected route | v2 route | v3b route | v3b success |
|
| 245 |
+
| --- | --- | --- | --- | --- |
|
| 246 |
+
| `text_genome_soul` | `text_v2` | `text_v2` | `text_v2` | yes |
|
| 247 |
+
| `text_fro` | `text_v2` | `text_v2` | `text_v2` | no |
|
| 248 |
+
| `code_fibonacci` | `code_v2` | `code_v2` | `code_v2` | yes |
|
| 249 |
+
| `code_prime` | `code_v2` | `text_v2` | `code_v2` | yes |
|
| 250 |
+
| `math_linear` | `math_v1` | `math_v1` | `math_v1` | yes |
|
| 251 |
+
| `math_speed` | `math_v1` | `math_v1` | `math_v1` | yes |
|
| 252 |
+
| `agentic_eval_plan` | `agentic_v1` | `agentic_v1` | `agentic_v1` | yes |
|
| 253 |
+
| `complex_multisoul` | `orchestrator_v1` | `orchestrator_v1` | `orchestrator_v1` | yes |
|
| 254 |
+
|
| 255 |
+
Interpretation:
|
| 256 |
+
|
| 257 |
+
```text
|
| 258 |
+
SwarmLM Orchestrator v3b restores full route accuracy on the controlled 8-task cascade suite and corrects the previous code_prime routing failure.
|
| 259 |
+
```
|
| 260 |
+
|
| 261 |
+
The remaining `text_fro` failure is not routing-related: it routes correctly to `text_v2`, but the generated specialist output does not match the current FRO-specific marker set. This result should therefore be reported as a controlled cascade-suite result, not as evidence of general routing robustness.
|
| 262 |
+
|
| 263 |
## What This Release Supports
|
| 264 |
|
| 265 |
This release supports:
|
|
|
|
| 298 |
|
| 299 |
> RTH-LM / SwarmLM v2 is a modular Genome/Soul research system. A shared frozen Genome supports multiple rank-512 specialist Souls, while `orchestrator_v2` routes controlled requests to specialist executors. In a controlled cascade evaluation, SwarmLM v2 reached 87.5% route accuracy and 75% end-to-end cascade success.
|
| 300 |
|
| 301 |
+
Updated Orchestrator v3b description:
|
| 302 |
+
|
| 303 |
+
> SwarmLM Orchestrator v3b improves centralized routing over v2 while preserving the same frozen Genome and v2 specialist Souls. On the controlled 8-task cascade suite, v3b reached 100% route accuracy and 87.5% cascade success, correcting the previous `code_prime` routing failure.
|
| 304 |
+
|
| 305 |
## License
|
| 306 |
|
| 307 |
Model artifacts are released for research and non-commercial use under the project license unless a separate commercial license is granted by RTH Italia.
|