DEVİM 336M Research
DEVİM is a Turkish-first neural language-model research program. This repository is the public research card for the controlled matched-scale ~336M probe.
No model weights are released here yet. The present plain causal Transformer is a controlled experimental substrate, not a permanent architectural commitment.
Controlled model
- Parameters: 336,390,144
- Layers: 24
d_model: 1024- Attention heads: 16
- Context length: 512
- Vocabulary: 32,768
- Training scope: natural-only Phase-P
Gate-B result
Gate-B training reached 36,755 optimizer steps and 1,200,012,349 supervised tokens.
Matched Gate-B measurements:
| Measurement | matched 110M | matched 336M |
|---|---|---|
| Breadth macro | 0.3364583 | 0.4333333 |
| FORM EOS | 0.8333333 | 0.96875 |
| FORM repetition | 0.1666667 | 0.0416667 |
| V1.8 macro | 0.44125 | 0.44125 |
| V1.8 competencies above chance | 4 | 4 |
| Positive-control macro | 0.7142857 | 0.8571429 |
Breadth delta was +0.096875, below the preregistered +0.20 requirement. The minimum 336M breadth-family result was 0.15, below the preregistered 0.45 requirement.
Formal decision:
GATE_B_CAPACITY_EFFECT_NOT_ESTABLISHED_OPEN_ARCHITECTURE_DESIGN_GATE
capacity_effect_supported: false
Interpretation
Scaling improved several readouts, especially FORM and positive controls, but did not establish the preregistered capacity effect. The evidence does not justify treating parameter count as the missing explanation for the target capability.
This is also not evidence that Transformers cannot learn the target capability. The result opens an architecture-design gate for controlled alternatives and matched ordinary baselines.
Release boundary
Not released here:
- model weights or checkpoints
- optimizer state
- hidden evaluator items
- rights-restricted corpus text
- private prompts or unpublished strategic mechanisms
Public research text is shared selectively. Do not describe this unreleased model as open-source or open-weight.
Project: https://devim.org
Source: https://devim.org
Research lead: Behram Bazo
Research text in this repository is intended as CC BY 4.0 unless stated otherwise. Future weights, code, datasets or evaluator artifacts may use separate licenses.