Sowrabhm commited on
Commit
b7f5c51
·
verified ·
1 Parent(s): 6b7ecc1

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +144 -0
README.md ADDED
@@ -0,0 +1,144 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-4.0
3
+ language:
4
+ - en
5
+ tags:
6
+ - onnx
7
+ - ner
8
+ - transaction-extraction
9
+ - sms-parsing
10
+ - gliner2
11
+ - deberta
12
+ - on-device
13
+ - mobile
14
+ library_name: onnxruntime
15
+ pipeline_tag: token-classification
16
+ ---
17
+
18
+ # Model Card: fintext-extractor
19
+
20
+ GLiNER2-based two-stage NER model that extracts structured transaction data from bank SMS and push notifications. Designed for on-device inference on mobile and desktop, with ONNX Runtime as the inference backend.
21
+
22
+ ## Architecture
23
+
24
+ fintext-extractor uses a **two-stage pipeline** to maximize both speed and accuracy:
25
+
26
+ 1. **Stage 1 -- Classification:** A DeBERTa-v3-large binary classifier determines whether an incoming message is a completed transaction (`is_transaction: yes/no`). Non-transaction messages (OTPs, promotional alerts, balance reminders) are filtered out early, keeping latency low.
27
+
28
+ 2. **Stage 2 -- Extraction:** A GLiNER2-large extraction model with a LoRA adapter runs only on messages classified as transactions. It extracts structured fields: amount, date, transaction type, description, and masked account digits.
29
+
30
+ This two-stage design means the heavier extraction model is invoked only when needed, reducing average inference cost on mixed message streams.
31
+
32
+ ## Extracted Fields
33
+
34
+ | Field | Type | Description |
35
+ |-------|------|-------------|
36
+ | `is_transaction` | bool | Whether the message is a completed transaction |
37
+ | `transaction_amount` | float | Numeric amount (e.g., 5000.00) |
38
+ | `transaction_type` | str | DEBIT or CREDIT |
39
+ | `transaction_date` | str | Date in DD-MM-YYYY format |
40
+ | `transaction_description` | str | Merchant or person name |
41
+ | `masked_account_digits` | str | Last 4 digits of card/account |
42
+
43
+ ## Model Files
44
+
45
+ | File | Size | Description |
46
+ |------|------|-------------|
47
+ | `onnx/deberta_classifier_fp16.onnx` + `.data` | ~830 MB | Classification model (FP16) |
48
+ | `onnx/deberta_classifier_fp32.onnx` + `.data` | ~1.66 GB | Classification model (FP32) |
49
+ | `onnx/extraction_full_fp16.onnx` + `.data` | ~930 MB | Extraction model (FP16) |
50
+ | `onnx/extraction_full_fp32.onnx` + `.data` | ~1.9 GB | Extraction model (FP32) |
51
+ | `tokenizer/` | ~11 MB | Classification tokenizer |
52
+ | `tokenizer_extraction/` | ~11 MB | Extraction tokenizer |
53
+
54
+ FP16 variants are recommended for most use cases. FP32 variants are provided for environments that do not support half-precision.
55
+
56
+ ## Quick Start (Python)
57
+
58
+ ```python
59
+ from fintext import FintextExtractor
60
+
61
+ extractor = FintextExtractor.from_pretrained("Sowrabhm/fintext-extractor")
62
+ result = extractor.extract("Rs.5,000 debited from a/c XX1234 for Amazon Pay on 08-Mar-26")
63
+ print(result)
64
+ # {'is_transaction': True, 'transaction_amount': 5000.0, 'transaction_type': 'DEBIT',
65
+ # 'transaction_date': '08-03-2026', 'transaction_description': 'Amazon Pay',
66
+ # 'masked_account_digits': '1234'}
67
+ ```
68
+
69
+ ## Direct ONNX Runtime Usage
70
+
71
+ If you prefer not to install the `fintext` library, you can run the ONNX models directly:
72
+
73
+ ```python
74
+ import numpy as np
75
+ import onnxruntime as ort
76
+ from tokenizers import Tokenizer
77
+
78
+ # Load classification model and tokenizer
79
+ cls_session = ort.InferenceSession("onnx/deberta_classifier_fp16.onnx")
80
+ tokenizer = Tokenizer.from_file("tokenizer/tokenizer.json")
81
+
82
+ # Tokenize input
83
+ text = "Rs.5,000 debited from a/c XX1234 for Amazon Pay on 08-Mar-26"
84
+ encoding = tokenizer.encode(text)
85
+ input_ids = np.array([encoding.ids], dtype=np.int64)
86
+ attention_mask = np.array([encoding.attention_mask], dtype=np.int64)
87
+
88
+ # Run classification
89
+ cls_output = cls_session.run(None, {
90
+ "input_ids": input_ids,
91
+ "attention_mask": attention_mask,
92
+ })
93
+ is_transaction = np.argmax(cls_output[0], axis=-1)[0] == 1
94
+
95
+ # If classified as a transaction, run extraction
96
+ if is_transaction:
97
+ ext_session = ort.InferenceSession("onnx/extraction_full_fp16.onnx")
98
+ ext_tokenizer = Tokenizer.from_file("tokenizer_extraction/tokenizer.json")
99
+ # ... tokenize and run extraction session
100
+ ```
101
+
102
+ ## Training
103
+
104
+ The models were fine-tuned from the following base checkpoints:
105
+
106
+ - **Classifier:** [microsoft/deberta-v3-large](https://huggingface.co/microsoft/deberta-v3-large) with LoRA (r=16, alpha=32)
107
+ - **Extractor:** [fastino/gliner2-large-v1](https://huggingface.co/fastino/gliner2-large-v1) with LoRA extraction adapter
108
+
109
+ Training used the GLiNER2 multi-task schema, combining binary classification (`is_transaction`) with structured extraction (`transaction_info`) in a single training loop. LoRA adapters keep the trainable parameter count low, enabling fine-tuning on consumer GPUs.
110
+
111
+ ## Metrics
112
+
113
+ | Metric | Value |
114
+ |--------|-------|
115
+ | Classification accuracy | 0.80 |
116
+ | Amount extraction accuracy | 1.00 |
117
+ | Type extraction accuracy | 1.00 |
118
+ | Digits extraction accuracy | 1.00 |
119
+ | Avg latency (FP16, CPU) | 47 ms |
120
+
121
+ Metrics were evaluated on a held-out test split. Latency measured on a single-threaded ONNX Runtime CPU session.
122
+
123
+ ## Limitations
124
+
125
+ - **Regional focus:** Primarily trained on Indian bank SMS formats (Rs., INR, currency symbols common in India). Performance on other regional formats has not been evaluated.
126
+ - **English only:** The model supports English language messages only.
127
+ - **Span extraction, not generation:** Field values must exist verbatim in the input text. The model extracts spans rather than generating new text.
128
+ - **Synthetic evaluation data:** The evaluation metrics above were computed on synthetic data. Real-world accuracy may differ.
129
+
130
+ ## Use Cases
131
+
132
+ - Personal finance apps
133
+ - Expense tracking and categorization
134
+ - Transaction monitoring and alerting
135
+ - Bank statement reconciliation from SMS/notifications
136
+
137
+ ## License
138
+
139
+ This model is released under the [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) license.
140
+
141
+ ## Links
142
+
143
+ - **GitHub:** [https://github.com/sowrabhmv/fintext-extractor](https://github.com/sowrabhmv/fintext-extractor)
144
+ - **Notebooks:** See the GitHub repo for cookbook examples and training notebooks