FluidAudio Web β raw model weights
Model weights for fluidaudio-web: fully in-browser inference (WebGPU + WebAssembly) for FluidAudio's speech models.
These are raw weights only β no ONNX, no onnxruntime. Each engine is a hand-written forward pass (WebGPU/WASM/JS) that loads these tensors directly and, where applicable, dequantizes them in-shader. Weights are extracted from the source models and parity-verified against the originals before publishing.
Layout (one folder per engine)
| folder | model | weights |
|---|---|---|
vad/ |
Silero VAD v5 (16 kHz) | fp32 |
parakeet/ |
Parakeet TDT 0.6B v3 | int8 encoder + fp32 decoder/joint |
nemotron/ |
Nemotron 3.5 streaming ASR 0.6B (40 langs) | int8 encoder + fp32 decoder |
eou/ |
Parakeet EOU 120M (end-of-utterance ASR) | fp16 encoder + fp32 decoder |
voicechat-stt/ |
NVIDIA VoiceChat-11B user-transcription chain (609M causal FastConformer + RNNT) | fp16 encoder + fp32 decoder |
whisper/ |
Whisper base | fp32 encoder + fp32 decoder |
sortformer/ |
Sortformer streaming diarization 4spk v2.1 | int8 encoder + fp32 head |
kokoro/ |
Kokoro TTS 82M β English | fp32 |
kokoro-zh/ |
Kokoro TTS 82M β Chinese | fp32 |
Format
*.binβ concatenated little-endian tensors.manifest.jsonβname -> { dims, offset, len }(offset/len in elements, not bytes).
Load: fetch the .bin, slice each tensor per the manifest. Extraction scripts and
the forward passes live in the fluidaudio-web repo (scripts/, src/engines/).
Quantized engines additionally carry per-tensor scales / palettes (documented per folder) for in-shader dequant.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support