Automatic Speech Recognition
MLX
Russian
English
gigaam
apple-silicon
russian
conformer
ctc
Eval Results (legacy)
Instructions to use aystream/GigaAM-v3-e2e-ctc-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use aystream/GigaAM-v3-e2e-ctc-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download aystream/GigaAM-v3-e2e-ctc-mlx --local-dir GigaAM-v3-e2e-ctc-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Upload folder using huggingface_hub
Browse files- README.md +73 -0
- config.json +32 -0
README.md
ADDED
|
@@ -0,0 +1,73 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: mlx
|
| 3 |
+
license: mit
|
| 4 |
+
language:
|
| 5 |
+
- ru
|
| 6 |
+
- en
|
| 7 |
+
tags:
|
| 8 |
+
- automatic-speech-recognition
|
| 9 |
+
- mlx
|
| 10 |
+
- apple-silicon
|
| 11 |
+
- russian
|
| 12 |
+
- gigaam
|
| 13 |
+
- conformer
|
| 14 |
+
- ctc
|
| 15 |
+
base_model: ai-sage/GigaAM-v3
|
| 16 |
+
pipeline_tag: automatic-speech-recognition
|
| 17 |
+
model-index:
|
| 18 |
+
- name: GigaAM-v3-e2e-ctc-mlx
|
| 19 |
+
results:
|
| 20 |
+
- task:
|
| 21 |
+
type: automatic-speech-recognition
|
| 22 |
+
metrics:
|
| 23 |
+
- name: RTF (M2 Max)
|
| 24 |
+
type: rtf
|
| 25 |
+
value: 0.006
|
| 26 |
+
---
|
| 27 |
+
|
| 28 |
+
# GigaAM v3 e2e CTC — MLX
|
| 29 |
+
|
| 30 |
+
MLX port of [GigaAM-v3](https://github.com/salute-developers/GigaAM) for fast Russian speech recognition on Apple Silicon. **180x realtime** on M2 Max.
|
| 31 |
+
|
| 32 |
+
## Usage
|
| 33 |
+
|
| 34 |
+
```bash
|
| 35 |
+
pip install gigaam-mlx
|
| 36 |
+
```
|
| 37 |
+
|
| 38 |
+
```python
|
| 39 |
+
from gigaam_mlx import load_model, transcribe
|
| 40 |
+
|
| 41 |
+
model, tokenizer = load_model() # downloads weights automatically
|
| 42 |
+
text = transcribe(model, tokenizer, "recording.wav")
|
| 43 |
+
print(text)
|
| 44 |
+
```
|
| 45 |
+
|
| 46 |
+
Or via CLI:
|
| 47 |
+
|
| 48 |
+
```bash
|
| 49 |
+
gigaam-mlx recording.wav
|
| 50 |
+
```
|
| 51 |
+
|
| 52 |
+
## Performance
|
| 53 |
+
|
| 54 |
+
MacBook Pro M2 Max, 20-second chunk:
|
| 55 |
+
|
| 56 |
+
| Backend | Time | Realtime |
|
| 57 |
+
|---|---|---|
|
| 58 |
+
| **MLX CTC (this)** | **0.11s** | **180x** |
|
| 59 |
+
| PyTorch MPS RNNT | 0.76s | 26x |
|
| 60 |
+
| ONNX CPU CTC | 1.66s | 12x |
|
| 61 |
+
|
| 62 |
+
## Model
|
| 63 |
+
|
| 64 |
+
- **Architecture:** Conformer (16 layers, 768d, 16 heads, RoPE) + CTC
|
| 65 |
+
- **Parameters:** 220M
|
| 66 |
+
- **Vocabulary:** 257 tokens (SentencePiece)
|
| 67 |
+
- **Features:** Punctuation, text normalization, Russian + English code-switching
|
| 68 |
+
|
| 69 |
+
## Links
|
| 70 |
+
|
| 71 |
+
- **Code:** [github.com/aystream/gigaam-mlx](https://github.com/aystream/gigaam-mlx)
|
| 72 |
+
- **Original:** [salute-developers/GigaAM](https://github.com/salute-developers/GigaAM) ([paper](https://arxiv.org/abs/2506.01192))
|
| 73 |
+
- **License:** MIT
|
config.json
ADDED
|
@@ -0,0 +1,32 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_type": "gigaam",
|
| 3 |
+
"model_variant": "v3_e2e_ctc",
|
| 4 |
+
"framework": "mlx",
|
| 5 |
+
"encoder": {
|
| 6 |
+
"feat_in": 64,
|
| 7 |
+
"n_layers": 16,
|
| 8 |
+
"d_model": 768,
|
| 9 |
+
"n_heads": 16,
|
| 10 |
+
"ff_expansion_factor": 4,
|
| 11 |
+
"conv_kernel_size": 5,
|
| 12 |
+
"subs_kernel_size": 5,
|
| 13 |
+
"subsampling": "conv1d",
|
| 14 |
+
"subsampling_factor": 4,
|
| 15 |
+
"self_attention_model": "rotary",
|
| 16 |
+
"rope_base": 5000
|
| 17 |
+
},
|
| 18 |
+
"head": {
|
| 19 |
+
"type": "ctc",
|
| 20 |
+
"num_classes": 257
|
| 21 |
+
},
|
| 22 |
+
"preprocessor": {
|
| 23 |
+
"sample_rate": 16000,
|
| 24 |
+
"n_mels": 64,
|
| 25 |
+
"hop_length": 160,
|
| 26 |
+
"win_length": 320,
|
| 27 |
+
"n_fft": 320,
|
| 28 |
+
"center": false
|
| 29 |
+
},
|
| 30 |
+
"tokenizer": "tokenizer.model",
|
| 31 |
+
"total_parameters": 220879361
|
| 32 |
+
}
|