EnCodec 32 kHz (decoder)
Meta's EnCodec at 32 kHz -- MusicGen's codec -- decode half, exported for loom.cpp. Family 11: codec tokens in, a waveform out.
This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.
Original model
Exported from facebook/encodec_32khz. Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.
License
cc-by-nc-4.0, inherited from the base model above.
Language(s)
(none tagged upstream)
a codec, not a language model: it carries no vocabulary and no language. The upstream repo carries NO license: tag; cc-by-nc-4.0 follows the MusicGen release this checkpoint was trained as part of, which is the stricter of the two readings. The EnCodec code itself is MIT, which covers the code and not these weights.
Usage
Run it with loom-py -- loom-py-rt on PyPI:
pip install -U "loom-py-rt[hub]"
import loom
model = loom.Model.from_pretrained("loom-ai-org/encodec-32khz-loom")
# The geometry a caller needs, declared by the file rather than looked up in a paper:
n_codebooks = model.hparam("codec.n_codebooks") # code streams per frame
codebook_size = model.hparam("codec.codebook_size") # valid id range per stream
frame_rate = model.hparam("codec.frame_rate", "f32") # codes per second
print(n_codebooks, codebook_size, frame_rate, model.contract["sample_rate"])
# Codes are FRAME-MAJOR: all `n_codebooks` codes for frame 0, then frame 1, and so on. This file
# is the DECODE half -- real codes come from the matching encoder, or an AR model that emits them.
frames = round(frame_rate) # one second of audio
codes = [[0] * n_codebooks for _ in range(frames)]
audio = model.codes2speech.infer(codes)
print(len(audio), "samples at", audio.sample_rate, "Hz =", round(audio.duration, 3), "s")
audio.save("out.wav")
# A flat list works too, and is what a driver that emitted the codes hands over. One that is not a
# whole number of frames is refused rather than reinterpreted at a different width.
audio = model.codes2speech.infer([0] * (frames * n_codebooks))
The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...)
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
model.driver_source prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See loom-py for the API and
loom.cpp for what the engine does between the two.
Known limitations
This is the DECODE half only. encode is audio-in/codes-out -- a different contract with a different modality pair -- and no model that decodes through this codec ever calls it, so exporting it would be weight in the file for a door nothing opens. To go the other way, use the upstream checkpoint.
It takes 4 code streams per frame at 50 frames per second, and one frame decodes to 640 samples. That is the 2.2 kbps bandwidth this checkpoint is configured at; codes from EnCodec at another bandwidth are a different number of streams and are refused on the width rather than decoded.
Its decoder contains an LSTM, which the engine walks per frame. The file carries three graph topologies rather than one, with the recurrence run between them -- the timestep loop is inside the engine, not in the script, and the whole thing is invisible from the outside. It is why a clip costs about 0.44x real time here where a purely convolutional codec is far cheaper: 8.8 ms per frame, flat from 25 frames to 400.
It does not undo a delay pattern. An AR model that emits these codes typically offsets stream k by k steps; realigning them is a property of that model, not of the codec, so feed it aligned codes.
Files
encodec-32khz.gguf-- the model, exported with loom-exporter.
- Downloads last month
- 27
We're not able to determine the quantization variants.
Model tree for loom-ai-org/encodec-32khz-loom
Base model
facebook/encodec_32khz