aystream commited on
Commit
fb0dfb7
·
verified ·
1 Parent(s): 208c2f8

Upload folder using huggingface_hub

Browse files
Files changed (2) hide show
  1. README.md +73 -0
  2. config.json +32 -0
README.md ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: mlx
3
+ license: mit
4
+ language:
5
+ - ru
6
+ - en
7
+ tags:
8
+ - automatic-speech-recognition
9
+ - mlx
10
+ - apple-silicon
11
+ - russian
12
+ - gigaam
13
+ - conformer
14
+ - ctc
15
+ base_model: ai-sage/GigaAM-v3
16
+ pipeline_tag: automatic-speech-recognition
17
+ model-index:
18
+ - name: GigaAM-v3-e2e-ctc-mlx
19
+ results:
20
+ - task:
21
+ type: automatic-speech-recognition
22
+ metrics:
23
+ - name: RTF (M2 Max)
24
+ type: rtf
25
+ value: 0.006
26
+ ---
27
+
28
+ # GigaAM v3 e2e CTC — MLX
29
+
30
+ MLX port of [GigaAM-v3](https://github.com/salute-developers/GigaAM) for fast Russian speech recognition on Apple Silicon. **180x realtime** on M2 Max.
31
+
32
+ ## Usage
33
+
34
+ ```bash
35
+ pip install gigaam-mlx
36
+ ```
37
+
38
+ ```python
39
+ from gigaam_mlx import load_model, transcribe
40
+
41
+ model, tokenizer = load_model() # downloads weights automatically
42
+ text = transcribe(model, tokenizer, "recording.wav")
43
+ print(text)
44
+ ```
45
+
46
+ Or via CLI:
47
+
48
+ ```bash
49
+ gigaam-mlx recording.wav
50
+ ```
51
+
52
+ ## Performance
53
+
54
+ MacBook Pro M2 Max, 20-second chunk:
55
+
56
+ | Backend | Time | Realtime |
57
+ |---|---|---|
58
+ | **MLX CTC (this)** | **0.11s** | **180x** |
59
+ | PyTorch MPS RNNT | 0.76s | 26x |
60
+ | ONNX CPU CTC | 1.66s | 12x |
61
+
62
+ ## Model
63
+
64
+ - **Architecture:** Conformer (16 layers, 768d, 16 heads, RoPE) + CTC
65
+ - **Parameters:** 220M
66
+ - **Vocabulary:** 257 tokens (SentencePiece)
67
+ - **Features:** Punctuation, text normalization, Russian + English code-switching
68
+
69
+ ## Links
70
+
71
+ - **Code:** [github.com/aystream/gigaam-mlx](https://github.com/aystream/gigaam-mlx)
72
+ - **Original:** [salute-developers/GigaAM](https://github.com/salute-developers/GigaAM) ([paper](https://arxiv.org/abs/2506.01192))
73
+ - **License:** MIT
config.json ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "gigaam",
3
+ "model_variant": "v3_e2e_ctc",
4
+ "framework": "mlx",
5
+ "encoder": {
6
+ "feat_in": 64,
7
+ "n_layers": 16,
8
+ "d_model": 768,
9
+ "n_heads": 16,
10
+ "ff_expansion_factor": 4,
11
+ "conv_kernel_size": 5,
12
+ "subs_kernel_size": 5,
13
+ "subsampling": "conv1d",
14
+ "subsampling_factor": 4,
15
+ "self_attention_model": "rotary",
16
+ "rope_base": 5000
17
+ },
18
+ "head": {
19
+ "type": "ctc",
20
+ "num_classes": 257
21
+ },
22
+ "preprocessor": {
23
+ "sample_rate": 16000,
24
+ "n_mels": 64,
25
+ "hop_length": 160,
26
+ "win_length": 320,
27
+ "n_fft": 320,
28
+ "center": false
29
+ },
30
+ "tokenizer": "tokenizer.model",
31
+ "total_parameters": 220879361
32
+ }