Instructions to use WaveCut/FLUX.2-klein-9B-OrbitQuant-W4A4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use WaveCut/FLUX.2-klein-9B-OrbitQuant-W4A4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("WaveCut/FLUX.2-klein-9B-OrbitQuant-W4A4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Add files using upload-large-folder tool
Browse files- .gitattributes +2 -0
- LICENSE.md +53 -0
- README.md +149 -0
- SHA256SUMS +21 -0
- assets/image_generation_comparison_matrix.webp +3 -0
- benchmark/build.json +10 -0
- benchmark/summary.json +111 -0
- model_index.json +26 -0
- orbitquant_pipeline_manifest.json +514 -0
- prompts.json +104 -0
- scheduler/scheduler_config.json +18 -0
- text_encoder/config.json +100 -0
- text_encoder/generation_config.json +7 -0
- text_encoder/model-00001-of-00002.safetensors +3 -0
- text_encoder/model-00002-of-00002.safetensors +3 -0
- text_encoder/model.safetensors.index.json +659 -0
- tokenizer/chat_template.jinja +89 -0
- tokenizer/tokenizer.json +3 -0
- tokenizer/tokenizer_config.json +15 -0
- transformer/config.json +53 -0
- transformer/diffusion_pytorch_model.safetensors +3 -0
- vae/config.json +41 -0
- vae/diffusion_pytorch_model.safetensors +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
assets/image_generation_comparison_matrix.webp filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
tokenizer/tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
LICENSE.md
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
FLUX Non-Commercial License v2.1
|
| 2 |
+
|
| 3 |
+
Black Forest Labs Inc. (“we” or “our” or “Company”) is pleased to make the weights, parameters, and inference code for the FLUX Model (as defined below) freely available for your non-commercial and non-production use as set forth in this FLUX Non-Commercial License (“License”). “FLUX Model” includes, individually and collectively, the models denoted as FLUX.x [dev], where “.x” denotes the FLUX Model version number, and models made available under the License, as indicated by a license notice that is included in or attached to the work, and their elements which includes algorithms, software, checkpoints, parameters, source code (inference code, evaluation code, and if applicable, fine-tuning code) and any other materials associated with the FLUX AI models made available by Company under this License, including if any, the technical documentation, manuals, and instructions for the use and operation thereof. Note that we may also make available certain elements of what is included in the definition of “FLUX Model” under a separate license, such as the inference code, and nothing in this License will be deemed to restrict or limit any other licenses granted by us in such elements.
|
| 4 |
+
|
| 5 |
+
By downloading, accessing, using, Distributing (as defined below), or creating a Derivative (as defined below) of the FLUX Model, you agree to the terms of this License. If you do not agree to this License, then you do not have any rights to access, use, Distribute or create a Derivative of the FLUX Model and you must immediately cease using the FLUX Model. If you are agreeing to be bound by the terms of this License on behalf of your employer or other entity, you represent and warrant to us that you have full legal authority to bind your employer or such entity to this License. If you do not have the requisite authority, you may not accept the License or access the FLUX Model on behalf of your employer or other entity.
|
| 6 |
+
|
| 7 |
+
1. Definitions.
|
| 8 |
+
- a. “Derivative” means any (i) modified version of the FLUX Model (including but not limited to any customized or fine-tuned version thereof), (ii) work based on the FLUX Model, or (iii) any other derivative work thereof. For the avoidance of doubt, Outputs are not considered Derivatives under this License.
|
| 9 |
+
- b. “Distribution” or “Distribute” or “Distributing” means providing or making available, by any means, a copy of the FLUX Model and/or the Derivatives as the case may be.
|
| 10 |
+
- c. “Non-Commercial Purpose” means any of the following uses, but only so far as you do not receive any direct or indirect payment arising from the use of the FLUX Model, Derivatives, or Content Filters (as defined below): (i) personal use for research, experimentation, and testing for the benefit of public knowledge, personal study, private entertainment, hobby projects, or otherwise not directly or indirectly connected to any commercial activities, business operations, or employment responsibilities; (ii) use by commercial or for-profit entities for testing, evaluation, or non-commercial research and development in a non-production environment; and (iii) use by any charitable organization for charitable purposes, or for testing or evaluation. For clarity, use (a) for revenue-generating activity, (b) in direct interactions with or that has impact on end users, or (c) to train, fine tune, or distill other models for commercial use, in each case, is not a Non-Commercial Purpose.
|
| 11 |
+
- d. “Outputs” means any content generated by the operation of the FLUX Model or Derivatives from an input (such as an image input) or prompt (i.e., text instructions) provided by users. For the avoidance of doubt, Outputs do not include any components of the FLUX Model, such as any fine-tuned versions of the FLUX Model, the weights, or parameters.
|
| 12 |
+
- e. “you” or “your” means the individual or entity entering into this License with Company.
|
| 13 |
+
|
| 14 |
+
2. License Grant.
|
| 15 |
+
- a. License. Subject to your compliance with this License, Company grants you a non-exclusive, worldwide, non-transferable, non-sublicensable, revocable, royalty free, and limited license to access, use, create Derivatives of, and Distribute the FLUX Model and Derivatives solely for your Non-Commercial Purposes. The foregoing license is personal to you, and you may not assign or sublicense this License or any other rights or obligations under this License without Company’s prior written consent; any such assignment or sublicense will be void and will automatically and immediately terminate this License. Any restrictions set forth herein regarding the FLUX Model also apply to any Derivative you create or that are created on your behalf.
|
| 16 |
+
- b. Non-Commercial Use Only. You may only access, use, Distribute, or create Derivatives of the FLUX Model or Derivatives for Non-Commercial Purposes. If you want to use a FLUX Model or a Derivative for any purpose that is not expressly authorized under this License, such as for a commercial activity, you must request a license from Company, which Company may grant to you in Company’s sole discretion and which additional use may be subject to a fee, royalty or other revenue share. Please see https://bfl.ai/licensing if you would like a commercial license.
|
| 17 |
+
- c. Reserved Rights. The grant of rights expressly set forth in this License are the complete grant of rights to use the FLUX Model, and no other licenses are granted, whether by waiver, estoppel, implication, equity, or otherwise. Company and its licensors reserve all rights not expressly granted by this License.
|
| 18 |
+
- d. Outputs. We claim no ownership rights in and to the Outputs. You are solely responsible for the Outputs you generate and their subsequent uses in accordance with this License. You may use Output for any purpose (including for commercial purposes), except as expressly prohibited herein. You may not use the Output to train, fine-tune, or distill a model that is competitive with a FLUX Model.
|
| 19 |
+
- e. You may access, use, Distribute, or create Output of the FLUX Model or Derivatives if you: (i) (A) implement and maintain content filtering measures (“Content Filters”) for your use of the FLUX Model or Derivatives to prevent the creation, display, transmission, generation, or dissemination of unlawful or infringing content, which may include Content Filters that we may make available for use with the FLUX Model (“Provided Content Filters”), or (B) ensure Output undergoes review for unlawful or infringing content before public or non-public distribution, display, transmission or dissemination; and (ii) ensure Output includes disclosure (or other indication) that the Output was generated or modified using artificial intelligence technologies to the extent required under applicable law.
|
| 20 |
+
|
| 21 |
+
3. Distribution. Subject to this License, you may Distribute copies of the FLUX Model and/or Derivatives made by you, under the following conditions:
|
| 22 |
+
- a. you must make available a copy of this License to third-party recipients of the FLUX Mode and/or Derivatives you Distribute, and specify that any rights to use the FLUX Model and/or Derivatives shall be directly granted by Company to said third-party recipients pursuant to this License;
|
| 23 |
+
- b. you must prominently display the following notice alongside the Distribution of the FLUX Model or Derivative (such as via a “Notice” text file distributed as part of such FLUX Model or Derivative) (the “Attribution Notice”):
|
| 24 |
+
|
| 25 |
+
> This FLUX Model is licensed by Black Forest Labs Inc. under the FLUX Non-Commercial License. Copyright Black Forest Labs Inc. IN NO EVENT SHALL BLACK FOREST LABS INC. BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH USE OF THIS MODEL.
|
| 26 |
+
|
| 27 |
+
- c. in the case of Distribution of Derivatives made by you: (i) you must also include in the Attribution Notice a statement that you have modified the applicable FLUX Model; (ii) any terms and conditions you impose on any third-party recipients relating to Derivatives made by or for you shall neither limit such third-party recipients’ use of the FLUX Model or any Derivatives made by or for Company in accordance with this License nor conflict with any of its terms and conditions and must include disclaimer of warranties and limitation of liability provisions that are at least as protective of Company as those set forth herein; and (iii) you must not misrepresent or imply, through any means, that the Derivatives made by or for you and/or any modified version of the FLUX Model you Distribute under your name and responsibility is an official product of the Company or has been endorsed, approved or validated by the Company, unless you are authorized by Company to do so in writing.
|
| 28 |
+
|
| 29 |
+
4. Restrictions. You will not, and will not permit, assist or cause any third party to
|
| 30 |
+
- a. use, modify, copy, reproduce, create Derivatives of, or Distribute the FLUX Model (or any Derivative thereof, or any data produced by the FLUX Model), in whole or in part, (i) for any commercial or production purposes, (ii) military purposes, (iii) purposes of surveillance, including any research or development relating to surveillance, (iv) biometric processing, (v) in any manner that infringes, misappropriates, or otherwise violates (or is likely to infringe, misappropriate, or otherwise violate) any third party’s legal rights, including rights of publicity or “digital replica” rights, (vi) in any unlawful, fraudulent, defamatory, or abusive activity, (vii) to generate unlawful content, including child sexual abuse material, or non-consensual intimate images; or (viii) in any manner that violates any applicable law and any privacy or security laws, rules, regulations, directives, or governmental requirements (including the General Data Privacy Regulation (Regulation (EU) 2016/679), the California Consumer Privacy Act, any and all laws governing the processing of biometric information, and the EU Artificial Intelligence Act (Regulation (EU) 2024/1689), as well as all amendments and successor laws to any of the foregoing);
|
| 31 |
+
- b. alter or remove copyright and other proprietary notices which appear on or in any portion of the FLUX Model;
|
| 32 |
+
- c. utilize any equipment, device, software, or other means to circumvent or remove any security or protection used by Company in connection with the FLUX Model, or to circumvent or remove any usage restrictions, or to enable functionality disabled by FLUX Model;
|
| 33 |
+
- d. offer or impose any terms on the FLUX Model that alter, restrict, or are inconsistent with the terms of this License;
|
| 34 |
+
- e. violate any applicable U.S. and non-U.S. export control and trade sanctions laws (“Export Laws”) in connection with your use or Distribution of any FLUX Model;
|
| 35 |
+
- f. directly or indirectly Distribute, export, or otherwise transfer FLUX Model (i) to any individual, entity, or country prohibited by Export Laws; (ii) to anyone on U.S. or non-U.S. government restricted parties lists; (iii) for any purpose prohibited by Export Laws, including nuclear, chemical or biological weapons, or missile technology applications; (iv) use or download FLUX Model if you or they are (a) located in a comprehensively sanctioned jurisdiction, (b) currently listed on any U.S. or non-U.S. restricted parties list, or (c) for any purpose prohibited by Export Laws; and (v) will not disguise your location through IP proxying or other methods.
|
| 36 |
+
|
| 37 |
+
5. DISCLAIMERS. THE FLUX MODEL AND PROVIDED CONTENT FILTERS ARE PROVIDED “AS IS” AND “WITH ALL FAULTS” WITH NO WARRANTY OF ANY KIND, EXPRESS OR IMPLIED. COMPANY EXPRESSLY DISCLAIMS ALL REPRESENTATIONS AND WARRANTIES, EXPRESS OR IMPLIED, WHETHER BY STATUTE, CUSTOM, USAGE OR OTHERWISE AS TO ANY MATTERS RELATED TO THE FLUX MODEL AND PROVIDED CONTENT FILTERS, INCLUDING BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE, SATISFACTORY QUALITY, OR NON-INFRINGEMENT. COMPANY MAKES NO WARRANTIES OR REPRESENTATIONS THAT THE FLUX MODEL AND PROVIDED CONTENT FILTERS WILL BE ERROR FREE OR FREE OF VIRUSES OR OTHER HARMFUL COMPONENTS, OR PRODUCE ANY PARTICULAR RESULTS.
|
| 38 |
+
|
| 39 |
+
6. LIMITATION OF LIABILITY. TO THE FULLEST EXTENT PERMITTED BY LAW, IN NO EVENT WILL COMPANY BE LIABLE TO YOU OR YOUR EMPLOYEES, AFFILIATES, USERS, OFFICERS OR DIRECTORS (A) UNDER ANY THEORY OF LIABILITY, WHETHER BASED IN CONTRACT, TORT, NEGLIGENCE, STRICT LIABILITY, WARRANTY, OR OTHERWISE UNDER THIS LICENSE, OR (B) FOR ANY INDIRECT, CONSEQUENTIAL, EXEMPLARY, INCIDENTAL, PUNITIVE OR SPECIAL DAMAGES OR LOST PROFITS, EVEN IF COMPANY HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. THE FLUX MODEL, ITS CONSTITUENT COMPONENTS, PROVIDED CONTENT FILTERS, AND ANY OUTPUT (COLLECTIVELY, “MODEL MATERIALS”) ARE NOT DESIGNED OR INTENDED FOR USE IN ANY APPLICATION OR SITUATION WHERE FAILURE OR FAULT OF THE MODEL MATERIALS COULD REASONABLY BE ANTICIPATED TO LEAD TO SERIOUS INJURY OF ANY PERSON, INCLUDING POTENTIAL DISCRIMINATION OR VIOLATION OF AN INDIVIDUAL’S PRIVACY RIGHTS, OR TO SEVERE PHYSICAL, PROPERTY, OR ENVIRONMENTAL DAMAGE (EACH, A “HIGH-RISK USE”). IF YOU ELECT TO USE ANY OF THE MODEL MATERIALS FOR A HIGH-RISK USE, YOU DO SO AT YOUR OWN RISK. YOU AGREE TO DESIGN AND IMPLEMENT APPROPRIATE DECISION-MAKING AND RISK-MITIGATION PROCEDURES AND POLICIES IN CONNECTION WITH A HIGH-RISK USE SUCH THAT EVEN IF THERE IS A FAILURE OR FAULT IN ANY OF THE MODEL MATERIALS, THE SAFETY OF PERSONS OR PROPERTY AFFECTED BY THE ACTIVITY STAYS AT A LEVEL THAT IS REASONABLE, APPROPRIATE, AND LAWFUL FOR THE FIELD OF THE HIGH-RISK USE.
|
| 40 |
+
|
| 41 |
+
7. INDEMNIFICATION. You will indemnify, defend and hold harmless Company and our subsidiaries and affiliates, and each of our respective shareholders, directors, officers, employees, agents, successors, and assigns (collectively, the “Company Parties”) from and against any losses, liabilities, damages, fines, penalties, and expenses (including reasonable attorneys’ fees) incurred by any Company Party in connection with any claim, demand, allegation, lawsuit, proceeding, or investigation (collectively, “Claims”) arising out of or related to (a) your access to or use of the FLUX Model (including in connection with any Output, results or data generated from such access or use, or from your access or use of any Content Filters), including any High-Risk Use; (b) your Content Filters, including your failure to implement any Content Filters where required by this License such as in Section 2(e); (c) your violation of this License; or (d) your violation, misappropriation or infringement of any rights of another (including intellectual property or other proprietary rights and privacy rights). You will promptly notify the Company Parties of any such Claims, and cooperate with Company Parties in defending such Claims. You will also grant the Company Parties sole control of the defense or settlement, at Company’s sole option, of any Claims. This indemnity is in addition to, and not in lieu of, any other indemnities or remedies set forth in a written agreement between you and Company or the other Company Parties.
|
| 42 |
+
|
| 43 |
+
8. Termination; Survival.
|
| 44 |
+
a. This License will automatically terminate upon any breach by you of the terms of this License.
|
| 45 |
+
b. We may terminate this License, in whole or in part, at any time upon notice (including electronic) to you.
|
| 46 |
+
c. If you initiate any legal action or proceedings against Company or any other entity (including a cross-claim or counterclaim in a lawsuit), alleging that the FLUX Model, any Derivative, or Provided Content Filters, or any part thereof, infringe upon intellectual property or other rights owned or licensable by you, then any licenses granted to you under this License will immediately terminate as of the date such legal action or claim is filed or initiated.
|
| 47 |
+
d. Upon termination of this License, you must cease all use, access or Distribution of the FLUX Model, any Derivatives, and any Provided Content Filters. The following sections survive termination of this License: 2(c), 2(d), 4-11.
|
| 48 |
+
|
| 49 |
+
9. Third Party Materials. The FLUX Model and Provided Content Filters may contain third-party software or other components (including free and open source software) (all of the foregoing, “Third Party Materials”), which are subject to the license terms of the respective third-party licensors. Your dealings or correspondence with third parties and your use of or interaction with any Third Party Materials are solely between you and the third party. Company does not control or endorse, and makes no representations or warranties regarding, any Third Party Materials, and your access to and use of such Third Party Materials are at your own risk.
|
| 50 |
+
|
| 51 |
+
10. Trademarks. You have not been granted any trademark license as part of this License and may not use any name, logo or trademark associated with Company without the prior written permission of Company, except to the extent necessary to make the reference required in the Attribution Notice as specified above or as is reasonably necessary in describing the FLUX Model and its creators.
|
| 52 |
+
|
| 53 |
+
11. General. This License will be governed and construed under the laws of the State of Delaware without regard to conflicts of law provisions. If any provision or part of a provision of this License is unlawful, void or unenforceable, that provision or part of the provision is deemed severed from this License, and will not affect the validity and enforceability of any remaining provisions. The failure of Company to exercise or enforce any right or provision of this License will not operate as a waiver of such right or provision. This License does not confer any third-party beneficiary rights upon any other person or entity. This License, together with the documentation, contains the entire understanding between you and Company regarding the subject matter of this License, and supersedes all other written or oral agreements and understandings between you and Company regarding such subject matter.
|
README.md
ADDED
|
@@ -0,0 +1,149 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: flux-non-commercial-license
|
| 4 |
+
license_link: https://huggingface.co/black-forest-labs/FLUX.2-klein-9B/blob/main/LICENSE.md
|
| 5 |
+
base_model: black-forest-labs/FLUX.2-klein-9B
|
| 6 |
+
pipeline_tag: text-to-image
|
| 7 |
+
library_name: diffusers
|
| 8 |
+
tags:
|
| 9 |
+
- diffusers
|
| 10 |
+
- flux
|
| 11 |
+
- orbitquant
|
| 12 |
+
- quantization
|
| 13 |
+
- text-to-image
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# FLUX.2 Klein 9B OrbitQuant W4A4
|
| 17 |
+
|
| 18 |
+
This is a complete Diffusers pipeline derived from
|
| 19 |
+
[`black-forest-labs/FLUX.2-klein-9B`](https://huggingface.co/black-forest-labs/FLUX.2-klein-9B)
|
| 20 |
+
at revision `92196c8e11f7b6cf2b7493e037d8c5345c559216`.
|
| 21 |
+
|
| 22 |
+
Both compute-heavy components are compressed:
|
| 23 |
+
|
| 24 |
+
- transformer: 144 OrbitQuant W4A4 projections plus 3 AdaLN INT4 weight-only projections;
|
| 25 |
+
- Qwen3 text encoder: 252 OrbitQuant W4A4 projections; `lm_head` remains BF16.
|
| 26 |
+
|
| 27 |
+
The text encoder is quantized to permit a controlled comparison with the published
|
| 28 |
+
[`FLUX.2-klein-9B-SDNQ-uint4-static`](https://huggingface.co/WaveCut/FLUX.2-klein-9B-SDNQ-uint4-static)
|
| 29 |
+
pipeline. This is an extension of OrbitQuant's architecture-independent adapter; the
|
| 30 |
+
OrbitQuant paper itself leaves text encoders in BF16.
|
| 31 |
+
|
| 32 |
+
## Install
|
| 33 |
+
|
| 34 |
+
```bash
|
| 35 |
+
pip install "orbitquant[kernels]>=0.2.2" "diffusers>=0.39" "transformers>=5.13" accelerate
|
| 36 |
+
```
|
| 37 |
+
|
| 38 |
+
OrbitQuant uses packed low-bit inference by default. CUDA first uses an importable
|
| 39 |
+
native package and otherwise uses the Triton packed fallback. It does not silently
|
| 40 |
+
materialize all weights in BF16. The native CUDA package can be built locally without
|
| 41 |
+
publishing to Kernel Hub; see the
|
| 42 |
+
[`OrbitQuant kernel instructions`](https://github.com/iamwavecut/OrbitQuant/blob/main/docs/kernel-audit.md#local-native-package).
|
| 43 |
+
|
| 44 |
+
## Diffusers
|
| 45 |
+
|
| 46 |
+
```python
|
| 47 |
+
import torch
|
| 48 |
+
import orbitquant # registers the quantizer
|
| 49 |
+
from diffusers import Flux2KleinPipeline
|
| 50 |
+
|
| 51 |
+
pipe = Flux2KleinPipeline.from_pretrained(
|
| 52 |
+
"WaveCut/FLUX.2-klein-9B-OrbitQuant-W4A4",
|
| 53 |
+
torch_dtype=torch.bfloat16,
|
| 54 |
+
).to("cuda")
|
| 55 |
+
|
| 56 |
+
image = pipe(
|
| 57 |
+
prompt="An intricate orbital conservatory above Earth, documentary realism",
|
| 58 |
+
height=1024,
|
| 59 |
+
width=1024,
|
| 60 |
+
num_inference_steps=4,
|
| 61 |
+
guidance_scale=1.0,
|
| 62 |
+
generator=torch.Generator(device="cuda").manual_seed(0),
|
| 63 |
+
).images[0]
|
| 64 |
+
image.save("orbitquant.png")
|
| 65 |
+
```
|
| 66 |
+
|
| 67 |
+
## Transformers Component
|
| 68 |
+
|
| 69 |
+
The quantized Qwen3 component can also be loaded through Transformers:
|
| 70 |
+
|
| 71 |
+
```python
|
| 72 |
+
import torch
|
| 73 |
+
import orbitquant
|
| 74 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 75 |
+
|
| 76 |
+
repo_id = "WaveCut/FLUX.2-klein-9B-OrbitQuant-W4A4"
|
| 77 |
+
tokenizer = AutoTokenizer.from_pretrained(repo_id, subfolder="tokenizer")
|
| 78 |
+
text_encoder = AutoModelForCausalLM.from_pretrained(
|
| 79 |
+
repo_id,
|
| 80 |
+
subfolder="text_encoder",
|
| 81 |
+
torch_dtype=torch.bfloat16,
|
| 82 |
+
).to("cuda")
|
| 83 |
+
```
|
| 84 |
+
|
| 85 |
+
## Quantization
|
| 86 |
+
|
| 87 |
+
| Setting | Value |
|
| 88 |
+
| --- | --- |
|
| 89 |
+
| Weight / activation bits | W4A4 |
|
| 90 |
+
| Rotation | RPBH, seed 0 |
|
| 91 |
+
| Block policy | Largest power-of-two divisor of the input dimension |
|
| 92 |
+
| Codebook | Lloyd-Max version 2 |
|
| 93 |
+
| Row norms | BF16 |
|
| 94 |
+
| AdaLN | INT4 RTN, group size 64, BF16 activations |
|
| 95 |
+
| Runtime | `auto_fused` |
|
| 96 |
+
| Calibration data | None |
|
| 97 |
+
|
| 98 |
+
The packed transformer and text-encoder weight payload is 10.67 GB; including the VAE,
|
| 99 |
+
the weight payload is 10.84 GB. The complete pipeline before this card asset is 10.85 GB.
|
| 100 |
+
|
| 101 |
+
## A40 Benchmark
|
| 102 |
+
|
| 103 |
+
All variants ran as separate processes on the same NVIDIA A40 with Torch 2.9.1+cu128,
|
| 104 |
+
CUDA 12.8, Diffusers 0.39.0, Transformers 5.13.0, BF16 arithmetic, no CPU offload,
|
| 105 |
+
1024x1024 output, four steps, guidance 1.0, seed 0, and the same ten prompts.
|
| 106 |
+
|
| 107 |
+
| Variant | Load | Cold image | Hot mean | Prompt encode median | NVML peak | CUDA peak |
|
| 108 |
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 109 |
+
| BF16 | 8.877 s | 4.597 s | 3.911 s | 0.110 s | 40.83 GB | 37.31 GB |
|
| 110 |
+
| SDNQ UINT4 | 37.642 s | 25.421 s | 3.546 s | 0.151 s | 17.55 GB | 14.86 GB |
|
| 111 |
+
| OrbitQuant W4A4, native CUDA | 25.229 s | 9.617 s | 7.418 s | 0.239 s | 17.66 GB | 13.96 GB |
|
| 112 |
+
|
| 113 |
+
The native run exercised all 396 OrbitQuant projections through
|
| 114 |
+
`native_packed_matmul`; activation quantization used Triton CUDA. OrbitQuant matched
|
| 115 |
+
SDNQ's memory class and used 56.7% less peak NVML memory than BF16, but it was 2.09x
|
| 116 |
+
slower than SDNQ in hot generation on this A40. This artifact does not claim an
|
| 117 |
+
end-to-end speedup. Full machine-readable results are in `benchmark/summary.json`.
|
| 118 |
+
|
| 119 |
+
## Visual Comparison
|
| 120 |
+
|
| 121 |
+
The matrix contains full-resolution output from ten difficult prompts: micro-detail,
|
| 122 |
+
counting, nested architecture, original mixed-media style, abstract materials, English
|
| 123 |
+
fine print, Russian typography, Japanese typography, Chinese typography and a dense
|
| 124 |
+
multi-subject panorama. Columns use identical prompts, settings and seed. The WebP was
|
| 125 |
+
encoded at quality 95.
|
| 126 |
+
|
| 127 |
+

|
| 128 |
+
|
| 129 |
+
All three variants remained coherent and detailed. OrbitQuant preserved material detail,
|
| 130 |
+
reflections and dense compositions without blank or noisy outputs. SDNQ reproduced the
|
| 131 |
+
small English specification text more accurately. Exact object counts and fine CJK text
|
| 132 |
+
were unreliable for every variant. This is a subjective paired inspection, not an
|
| 133 |
+
objective quality metric.
|
| 134 |
+
|
| 135 |
+
## Limitations
|
| 136 |
+
|
| 137 |
+
- Optimized native CUDA measurements require a locally built matching ABI3 kernel package.
|
| 138 |
+
- Triton packed matmul is a compatible fallback, not the fastest measured backend.
|
| 139 |
+
- W4A4 activation quantization can move the denoising trajectory farther from BF16 than
|
| 140 |
+
weight-only UINT4.
|
| 141 |
+
- Quantizing the text encoder is outside the paper's default layer policy.
|
| 142 |
+
- This model inherits the FLUX Non-Commercial License from the source checkpoint.
|
| 143 |
+
|
| 144 |
+
## References
|
| 145 |
+
|
| 146 |
+
- [OrbitQuant paper](https://arxiv.org/abs/2607.02461)
|
| 147 |
+
- [OrbitQuant library](https://github.com/iamwavecut/OrbitQuant)
|
| 148 |
+
- [Source FLUX.2 Klein 9B model](https://huggingface.co/black-forest-labs/FLUX.2-klein-9B)
|
| 149 |
+
- [Controlled SDNQ comparison checkpoint](https://huggingface.co/WaveCut/FLUX.2-klein-9B-SDNQ-uint4-static)
|
SHA256SUMS
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
468d9f4332c0c895e9132c035982a40a12750092603a2e910913cb49aa887d3b ./LICENSE.md
|
| 2 |
+
7d4f666483bfadb61d7a7cbb504441b6945f3b03744770f1382b42b4512882e5 ./README.md
|
| 3 |
+
02f199e9e833501e8ee57418016ca3514376a90a6f952c46e2542c6d1ec3ff27 ./assets/image_generation_comparison_matrix.webp
|
| 4 |
+
8155fb3d937b8dca2113f5bd16dbba847104ed3420c34bdedc551b3e6f792120 ./benchmark/build.json
|
| 5 |
+
1eccd6b514c82c9bb71d7705ca92761f1b54e9daae881d010bad4ba0251254be ./benchmark/summary.json
|
| 6 |
+
bfbb65f55d0369e70d88eb612cf7fc233ee0296860f501e3493a4d67fe6117e4 ./model_index.json
|
| 7 |
+
78f90f2177f374ec3075831620e7a98fcddf853043082006738cb7b4ee057dd6 ./orbitquant_pipeline_manifest.json
|
| 8 |
+
04cdd53e5805e90a9db7e50e7d37f1018c0364d6351bb89cc7ce8e5416493c51 ./prompts.json
|
| 9 |
+
5d5af7e00dad78642b2bd51ddad2ba0f2758dbef9fb441a3dbee5ceb2fbcb8e3 ./scheduler/scheduler_config.json
|
| 10 |
+
b282b774e20c5f81feb2f1e1ce4a23d94cc4f7faa3dbd656ecd22af41c5dff12 ./text_encoder/config.json
|
| 11 |
+
a3d75a906863147b046df4cd04337b6bc2bcc66ef8ebf678511e909548c5f99d ./text_encoder/generation_config.json
|
| 12 |
+
baff83ea9d23ff652ef1b0647e7dfd0af153797183a6edf9098740b168067253 ./text_encoder/model-00001-of-00002.safetensors
|
| 13 |
+
147e9d7ecd892c90d5fecedfa0a3e0706b5bd234b2d27a8071fcc1508c29b834 ./text_encoder/model-00002-of-00002.safetensors
|
| 14 |
+
d494d73e77fbb3da94ad2c10ea08a24135adcf29ea473a25b455afceecd48ae3 ./text_encoder/model.safetensors.index.json
|
| 15 |
+
a55ee1b1660128b7098723e0abcd92caa0788061051c62d51cbe87d9cf1974d8 ./tokenizer/chat_template.jinja
|
| 16 |
+
be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506 ./tokenizer/tokenizer.json
|
| 17 |
+
1689852cc9c45010de040c8302a8acdc0d2c4c6c740dd7e9dd0a8c704e16eada ./tokenizer/tokenizer_config.json
|
| 18 |
+
fc2ecb80ad1d98f266413b0f326022a29633232ca4a75c7306309d7edb46de88 ./transformer/config.json
|
| 19 |
+
c5eaeecc5c474756a28de32d67edd8320ac44fcdb9fc5f71d83ab4cdeccac84e ./transformer/diffusion_pytorch_model.safetensors
|
| 20 |
+
cbbc4d8a187f8b9cc8adaef826e97e3ff588d4b154b3d63a99f098c3d2d39d8e ./vae/config.json
|
| 21 |
+
ca70d2202afe6415bdbcb8793ba8cd99fd159cfe6192381504d6c4d3036e0f04 ./vae/diffusion_pytorch_model.safetensors
|
assets/image_generation_comparison_matrix.webp
ADDED
|
Git LFS Details
|
benchmark/build.json
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"source_load_seconds": 17.03382627852261,
|
| 3 |
+
"save_seconds": 24.1477939337492,
|
| 4 |
+
"total_seconds": 67.16633238084614,
|
| 5 |
+
"process_rss_bytes": 12405612544,
|
| 6 |
+
"process_peak_rss_bytes": 12405612544,
|
| 7 |
+
"torch_version": "2.9.1+cu128",
|
| 8 |
+
"cuda_version": "12.8",
|
| 9 |
+
"gpu": "NVIDIA A40"
|
| 10 |
+
}
|
benchmark/summary.json
ADDED
|
@@ -0,0 +1,111 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"format": "flux2-klein-9b-quantizer-comparison-v1",
|
| 3 |
+
"source_model_id": "black-forest-labs/FLUX.2-klein-9B",
|
| 4 |
+
"source_revision": "92196c8e11f7b6cf2b7493e037d8c5345c559216",
|
| 5 |
+
"sdnq_model_id": "WaveCut/FLUX.2-klein-9B-SDNQ-uint4-static",
|
| 6 |
+
"sdnq_revision": "ed71b3f19ce640e88b66a2a743aabb8a613adeac",
|
| 7 |
+
"prompt_pack": "flux2_klein_9b_quantizer_stress_v1",
|
| 8 |
+
"prompt_count": 10,
|
| 9 |
+
"settings": {
|
| 10 |
+
"width": 1024,
|
| 11 |
+
"height": 1024,
|
| 12 |
+
"steps": 4,
|
| 13 |
+
"guidance": 1.0,
|
| 14 |
+
"seed": 0,
|
| 15 |
+
"dtype": "bfloat16",
|
| 16 |
+
"batch_size": 1,
|
| 17 |
+
"cpu_offload": false
|
| 18 |
+
},
|
| 19 |
+
"environment": {
|
| 20 |
+
"gpu": "NVIDIA A40",
|
| 21 |
+
"gpu_memory_bytes": 48305799168,
|
| 22 |
+
"torch": "2.9.1+cu128",
|
| 23 |
+
"cuda": "12.8",
|
| 24 |
+
"diffusers": "0.39.0",
|
| 25 |
+
"transformers": "5.13.0",
|
| 26 |
+
"orbitquant": "0.2.2",
|
| 27 |
+
"sdnq": "0.1.8"
|
| 28 |
+
},
|
| 29 |
+
"weight_payload_bytes": {
|
| 30 |
+
"bf16": 34706828582,
|
| 31 |
+
"sdnq_uint4": 12181412678,
|
| 32 |
+
"orbitquant_w4a4": 10838585126
|
| 33 |
+
},
|
| 34 |
+
"variants": {
|
| 35 |
+
"bf16": {
|
| 36 |
+
"load_seconds": 8.876504,
|
| 37 |
+
"cold_generation_seconds": 4.596600,
|
| 38 |
+
"hot_generation_mean_seconds": 3.910797,
|
| 39 |
+
"hot_generation_median_seconds": 3.911681,
|
| 40 |
+
"hot_generation_p95_seconds": 3.931342,
|
| 41 |
+
"prompt_encode_mean_seconds": 0.110661,
|
| 42 |
+
"prompt_encode_median_seconds": 0.110441,
|
| 43 |
+
"nvml_peak_bytes": 40825454592,
|
| 44 |
+
"cuda_peak_allocated_bytes": 37306757120,
|
| 45 |
+
"gpu_utilization_mean_percent": 96.865949,
|
| 46 |
+
"power_mean_watts": 279.082023,
|
| 47 |
+
"power_peak_watts": 307.282,
|
| 48 |
+
"energy_watt_hours_for_10_images": 3.035286
|
| 49 |
+
},
|
| 50 |
+
"sdnq_uint4": {
|
| 51 |
+
"load_seconds": 37.641839,
|
| 52 |
+
"cold_generation_seconds": 25.420842,
|
| 53 |
+
"hot_generation_mean_seconds": 3.546106,
|
| 54 |
+
"hot_generation_median_seconds": 3.545387,
|
| 55 |
+
"hot_generation_p95_seconds": 3.576177,
|
| 56 |
+
"prompt_encode_mean_seconds": 0.228912,
|
| 57 |
+
"prompt_encode_median_seconds": 0.151227,
|
| 58 |
+
"nvml_peak_bytes": 17547067392,
|
| 59 |
+
"cuda_peak_allocated_bytes": 14856482816,
|
| 60 |
+
"gpu_utilization_mean_percent": 97.187421,
|
| 61 |
+
"power_mean_watts": 274.029520,
|
| 62 |
+
"power_peak_watts": 302.010,
|
| 63 |
+
"energy_watt_hours_for_10_images": 2.704396
|
| 64 |
+
},
|
| 65 |
+
"orbitquant_w4a4_native": {
|
| 66 |
+
"load_seconds": 25.229073,
|
| 67 |
+
"cold_generation_seconds": 9.617251,
|
| 68 |
+
"hot_generation_mean_seconds": 7.417985,
|
| 69 |
+
"hot_generation_median_seconds": 7.414725,
|
| 70 |
+
"hot_generation_p95_seconds": 7.502579,
|
| 71 |
+
"prompt_encode_mean_seconds": 0.239160,
|
| 72 |
+
"prompt_encode_median_seconds": 0.239016,
|
| 73 |
+
"nvml_peak_bytes": 17660313600,
|
| 74 |
+
"cuda_peak_allocated_bytes": 13962950656,
|
| 75 |
+
"gpu_utilization_mean_percent": 99.561541,
|
| 76 |
+
"power_mean_watts": 289.659992,
|
| 77 |
+
"power_peak_watts": 306.366,
|
| 78 |
+
"energy_watt_hours_for_10_images": 5.979187,
|
| 79 |
+
"orbitquant_linear_count": 396,
|
| 80 |
+
"effective_runtime_mode": "native_packed_matmul",
|
| 81 |
+
"activation_backend": "triton_cuda"
|
| 82 |
+
},
|
| 83 |
+
"orbitquant_w4a4_triton_fallback": {
|
| 84 |
+
"load_seconds": 19.324567,
|
| 85 |
+
"cold_generation_seconds": 30.720916,
|
| 86 |
+
"hot_generation_mean_seconds": 19.308872,
|
| 87 |
+
"hot_generation_median_seconds": 19.307104,
|
| 88 |
+
"hot_generation_p95_seconds": 19.365284,
|
| 89 |
+
"prompt_encode_mean_seconds": 0.588539,
|
| 90 |
+
"prompt_encode_median_seconds": 0.588336,
|
| 91 |
+
"nvml_peak_bytes": 17936416768,
|
| 92 |
+
"cuda_peak_allocated_bytes": 14338371072,
|
| 93 |
+
"gpu_utilization_mean_percent": 99.818867,
|
| 94 |
+
"power_mean_watts": 287.027361,
|
| 95 |
+
"power_peak_watts": 302.896,
|
| 96 |
+
"energy_watt_hours_for_10_images": 15.405621,
|
| 97 |
+
"orbitquant_linear_count": 396,
|
| 98 |
+
"effective_runtime_mode": "triton_packed_matmul",
|
| 99 |
+
"activation_backend": "triton_cuda"
|
| 100 |
+
}
|
| 101 |
+
},
|
| 102 |
+
"comparison_asset": {
|
| 103 |
+
"path": "assets/image_generation_comparison_matrix.webp",
|
| 104 |
+
"format": "webp",
|
| 105 |
+
"quality": 95,
|
| 106 |
+
"method": 6,
|
| 107 |
+
"width": 3072,
|
| 108 |
+
"height": 10890,
|
| 109 |
+
"sha256": "02f199e9e833501e8ee57418016ca3514376a90a6f952c46e2542c6d1ec3ff27"
|
| 110 |
+
}
|
| 111 |
+
}
|
model_index.json
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_class_name": "Flux2KleinPipeline",
|
| 3 |
+
"_diffusers_version": "0.39.0",
|
| 4 |
+
"_name_or_path": "black-forest-labs/FLUX.2-klein-9B",
|
| 5 |
+
"is_distilled": true,
|
| 6 |
+
"scheduler": [
|
| 7 |
+
"diffusers",
|
| 8 |
+
"FlowMatchEulerDiscreteScheduler"
|
| 9 |
+
],
|
| 10 |
+
"text_encoder": [
|
| 11 |
+
"transformers",
|
| 12 |
+
"Qwen3ForCausalLM"
|
| 13 |
+
],
|
| 14 |
+
"tokenizer": [
|
| 15 |
+
"transformers",
|
| 16 |
+
"Qwen2Tokenizer"
|
| 17 |
+
],
|
| 18 |
+
"transformer": [
|
| 19 |
+
"diffusers",
|
| 20 |
+
"Flux2Transformer2DModel"
|
| 21 |
+
],
|
| 22 |
+
"vae": [
|
| 23 |
+
"diffusers",
|
| 24 |
+
"AutoencoderKLFlux2"
|
| 25 |
+
]
|
| 26 |
+
}
|
orbitquant_pipeline_manifest.json
ADDED
|
@@ -0,0 +1,514 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"format": "orbitquant-full-pipeline-v1",
|
| 3 |
+
"source_model_id": "black-forest-labs/FLUX.2-klein-9B",
|
| 4 |
+
"source_revision": "92196c8e11f7b6cf2b7493e037d8c5345c559216",
|
| 5 |
+
"source_license": "flux-non-commercial-license",
|
| 6 |
+
"quant_method": "orbitquant",
|
| 7 |
+
"weight_bits": 4,
|
| 8 |
+
"activation_bits": 4,
|
| 9 |
+
"runtime_mode": "auto_fused",
|
| 10 |
+
"activation_kernel_backend": "triton_cuda",
|
| 11 |
+
"rotation": "rpbh",
|
| 12 |
+
"rotation_seed": 0,
|
| 13 |
+
"codebook": "lloyd_max",
|
| 14 |
+
"codebook_version": 2,
|
| 15 |
+
"prompt_pack": "flux2_klein_9b_quantizer_stress_v1",
|
| 16 |
+
"quantized_components": [
|
| 17 |
+
"transformer",
|
| 18 |
+
"text_encoder"
|
| 19 |
+
],
|
| 20 |
+
"paper_policy_deviation": "The text encoder is quantized for parity with the SDNQ deployment checkpoint; the OrbitQuant paper leaves text encoders in BF16.",
|
| 21 |
+
"components": {
|
| 22 |
+
"transformer": {
|
| 23 |
+
"elapsed_seconds": 11.634634396061301,
|
| 24 |
+
"quantized_module_count": 144,
|
| 25 |
+
"adaln_module_count": 3,
|
| 26 |
+
"skipped_module_count": 6,
|
| 27 |
+
"quantized_modules": [
|
| 28 |
+
"transformer_blocks.0.attn.to_q",
|
| 29 |
+
"transformer_blocks.0.attn.to_k",
|
| 30 |
+
"transformer_blocks.0.attn.to_v",
|
| 31 |
+
"transformer_blocks.0.attn.to_out.0",
|
| 32 |
+
"transformer_blocks.0.attn.add_q_proj",
|
| 33 |
+
"transformer_blocks.0.attn.add_k_proj",
|
| 34 |
+
"transformer_blocks.0.attn.add_v_proj",
|
| 35 |
+
"transformer_blocks.0.attn.to_add_out",
|
| 36 |
+
"transformer_blocks.0.ff.linear_in",
|
| 37 |
+
"transformer_blocks.0.ff.linear_out",
|
| 38 |
+
"transformer_blocks.0.ff_context.linear_in",
|
| 39 |
+
"transformer_blocks.0.ff_context.linear_out",
|
| 40 |
+
"transformer_blocks.1.attn.to_q",
|
| 41 |
+
"transformer_blocks.1.attn.to_k",
|
| 42 |
+
"transformer_blocks.1.attn.to_v",
|
| 43 |
+
"transformer_blocks.1.attn.to_out.0",
|
| 44 |
+
"transformer_blocks.1.attn.add_q_proj",
|
| 45 |
+
"transformer_blocks.1.attn.add_k_proj",
|
| 46 |
+
"transformer_blocks.1.attn.add_v_proj",
|
| 47 |
+
"transformer_blocks.1.attn.to_add_out",
|
| 48 |
+
"transformer_blocks.1.ff.linear_in",
|
| 49 |
+
"transformer_blocks.1.ff.linear_out",
|
| 50 |
+
"transformer_blocks.1.ff_context.linear_in",
|
| 51 |
+
"transformer_blocks.1.ff_context.linear_out",
|
| 52 |
+
"transformer_blocks.2.attn.to_q",
|
| 53 |
+
"transformer_blocks.2.attn.to_k",
|
| 54 |
+
"transformer_blocks.2.attn.to_v",
|
| 55 |
+
"transformer_blocks.2.attn.to_out.0",
|
| 56 |
+
"transformer_blocks.2.attn.add_q_proj",
|
| 57 |
+
"transformer_blocks.2.attn.add_k_proj",
|
| 58 |
+
"transformer_blocks.2.attn.add_v_proj",
|
| 59 |
+
"transformer_blocks.2.attn.to_add_out",
|
| 60 |
+
"transformer_blocks.2.ff.linear_in",
|
| 61 |
+
"transformer_blocks.2.ff.linear_out",
|
| 62 |
+
"transformer_blocks.2.ff_context.linear_in",
|
| 63 |
+
"transformer_blocks.2.ff_context.linear_out",
|
| 64 |
+
"transformer_blocks.3.attn.to_q",
|
| 65 |
+
"transformer_blocks.3.attn.to_k",
|
| 66 |
+
"transformer_blocks.3.attn.to_v",
|
| 67 |
+
"transformer_blocks.3.attn.to_out.0",
|
| 68 |
+
"transformer_blocks.3.attn.add_q_proj",
|
| 69 |
+
"transformer_blocks.3.attn.add_k_proj",
|
| 70 |
+
"transformer_blocks.3.attn.add_v_proj",
|
| 71 |
+
"transformer_blocks.3.attn.to_add_out",
|
| 72 |
+
"transformer_blocks.3.ff.linear_in",
|
| 73 |
+
"transformer_blocks.3.ff.linear_out",
|
| 74 |
+
"transformer_blocks.3.ff_context.linear_in",
|
| 75 |
+
"transformer_blocks.3.ff_context.linear_out",
|
| 76 |
+
"transformer_blocks.4.attn.to_q",
|
| 77 |
+
"transformer_blocks.4.attn.to_k",
|
| 78 |
+
"transformer_blocks.4.attn.to_v",
|
| 79 |
+
"transformer_blocks.4.attn.to_out.0",
|
| 80 |
+
"transformer_blocks.4.attn.add_q_proj",
|
| 81 |
+
"transformer_blocks.4.attn.add_k_proj",
|
| 82 |
+
"transformer_blocks.4.attn.add_v_proj",
|
| 83 |
+
"transformer_blocks.4.attn.to_add_out",
|
| 84 |
+
"transformer_blocks.4.ff.linear_in",
|
| 85 |
+
"transformer_blocks.4.ff.linear_out",
|
| 86 |
+
"transformer_blocks.4.ff_context.linear_in",
|
| 87 |
+
"transformer_blocks.4.ff_context.linear_out",
|
| 88 |
+
"transformer_blocks.5.attn.to_q",
|
| 89 |
+
"transformer_blocks.5.attn.to_k",
|
| 90 |
+
"transformer_blocks.5.attn.to_v",
|
| 91 |
+
"transformer_blocks.5.attn.to_out.0",
|
| 92 |
+
"transformer_blocks.5.attn.add_q_proj",
|
| 93 |
+
"transformer_blocks.5.attn.add_k_proj",
|
| 94 |
+
"transformer_blocks.5.attn.add_v_proj",
|
| 95 |
+
"transformer_blocks.5.attn.to_add_out",
|
| 96 |
+
"transformer_blocks.5.ff.linear_in",
|
| 97 |
+
"transformer_blocks.5.ff.linear_out",
|
| 98 |
+
"transformer_blocks.5.ff_context.linear_in",
|
| 99 |
+
"transformer_blocks.5.ff_context.linear_out",
|
| 100 |
+
"transformer_blocks.6.attn.to_q",
|
| 101 |
+
"transformer_blocks.6.attn.to_k",
|
| 102 |
+
"transformer_blocks.6.attn.to_v",
|
| 103 |
+
"transformer_blocks.6.attn.to_out.0",
|
| 104 |
+
"transformer_blocks.6.attn.add_q_proj",
|
| 105 |
+
"transformer_blocks.6.attn.add_k_proj",
|
| 106 |
+
"transformer_blocks.6.attn.add_v_proj",
|
| 107 |
+
"transformer_blocks.6.attn.to_add_out",
|
| 108 |
+
"transformer_blocks.6.ff.linear_in",
|
| 109 |
+
"transformer_blocks.6.ff.linear_out",
|
| 110 |
+
"transformer_blocks.6.ff_context.linear_in",
|
| 111 |
+
"transformer_blocks.6.ff_context.linear_out",
|
| 112 |
+
"transformer_blocks.7.attn.to_q",
|
| 113 |
+
"transformer_blocks.7.attn.to_k",
|
| 114 |
+
"transformer_blocks.7.attn.to_v",
|
| 115 |
+
"transformer_blocks.7.attn.to_out.0",
|
| 116 |
+
"transformer_blocks.7.attn.add_q_proj",
|
| 117 |
+
"transformer_blocks.7.attn.add_k_proj",
|
| 118 |
+
"transformer_blocks.7.attn.add_v_proj",
|
| 119 |
+
"transformer_blocks.7.attn.to_add_out",
|
| 120 |
+
"transformer_blocks.7.ff.linear_in",
|
| 121 |
+
"transformer_blocks.7.ff.linear_out",
|
| 122 |
+
"transformer_blocks.7.ff_context.linear_in",
|
| 123 |
+
"transformer_blocks.7.ff_context.linear_out",
|
| 124 |
+
"single_transformer_blocks.0.attn.to_qkv_mlp_proj",
|
| 125 |
+
"single_transformer_blocks.0.attn.to_out",
|
| 126 |
+
"single_transformer_blocks.1.attn.to_qkv_mlp_proj",
|
| 127 |
+
"single_transformer_blocks.1.attn.to_out",
|
| 128 |
+
"single_transformer_blocks.2.attn.to_qkv_mlp_proj",
|
| 129 |
+
"single_transformer_blocks.2.attn.to_out",
|
| 130 |
+
"single_transformer_blocks.3.attn.to_qkv_mlp_proj",
|
| 131 |
+
"single_transformer_blocks.3.attn.to_out",
|
| 132 |
+
"single_transformer_blocks.4.attn.to_qkv_mlp_proj",
|
| 133 |
+
"single_transformer_blocks.4.attn.to_out",
|
| 134 |
+
"single_transformer_blocks.5.attn.to_qkv_mlp_proj",
|
| 135 |
+
"single_transformer_blocks.5.attn.to_out",
|
| 136 |
+
"single_transformer_blocks.6.attn.to_qkv_mlp_proj",
|
| 137 |
+
"single_transformer_blocks.6.attn.to_out",
|
| 138 |
+
"single_transformer_blocks.7.attn.to_qkv_mlp_proj",
|
| 139 |
+
"single_transformer_blocks.7.attn.to_out",
|
| 140 |
+
"single_transformer_blocks.8.attn.to_qkv_mlp_proj",
|
| 141 |
+
"single_transformer_blocks.8.attn.to_out",
|
| 142 |
+
"single_transformer_blocks.9.attn.to_qkv_mlp_proj",
|
| 143 |
+
"single_transformer_blocks.9.attn.to_out",
|
| 144 |
+
"single_transformer_blocks.10.attn.to_qkv_mlp_proj",
|
| 145 |
+
"single_transformer_blocks.10.attn.to_out",
|
| 146 |
+
"single_transformer_blocks.11.attn.to_qkv_mlp_proj",
|
| 147 |
+
"single_transformer_blocks.11.attn.to_out",
|
| 148 |
+
"single_transformer_blocks.12.attn.to_qkv_mlp_proj",
|
| 149 |
+
"single_transformer_blocks.12.attn.to_out",
|
| 150 |
+
"single_transformer_blocks.13.attn.to_qkv_mlp_proj",
|
| 151 |
+
"single_transformer_blocks.13.attn.to_out",
|
| 152 |
+
"single_transformer_blocks.14.attn.to_qkv_mlp_proj",
|
| 153 |
+
"single_transformer_blocks.14.attn.to_out",
|
| 154 |
+
"single_transformer_blocks.15.attn.to_qkv_mlp_proj",
|
| 155 |
+
"single_transformer_blocks.15.attn.to_out",
|
| 156 |
+
"single_transformer_blocks.16.attn.to_qkv_mlp_proj",
|
| 157 |
+
"single_transformer_blocks.16.attn.to_out",
|
| 158 |
+
"single_transformer_blocks.17.attn.to_qkv_mlp_proj",
|
| 159 |
+
"single_transformer_blocks.17.attn.to_out",
|
| 160 |
+
"single_transformer_blocks.18.attn.to_qkv_mlp_proj",
|
| 161 |
+
"single_transformer_blocks.18.attn.to_out",
|
| 162 |
+
"single_transformer_blocks.19.attn.to_qkv_mlp_proj",
|
| 163 |
+
"single_transformer_blocks.19.attn.to_out",
|
| 164 |
+
"single_transformer_blocks.20.attn.to_qkv_mlp_proj",
|
| 165 |
+
"single_transformer_blocks.20.attn.to_out",
|
| 166 |
+
"single_transformer_blocks.21.attn.to_qkv_mlp_proj",
|
| 167 |
+
"single_transformer_blocks.21.attn.to_out",
|
| 168 |
+
"single_transformer_blocks.22.attn.to_qkv_mlp_proj",
|
| 169 |
+
"single_transformer_blocks.22.attn.to_out",
|
| 170 |
+
"single_transformer_blocks.23.attn.to_qkv_mlp_proj",
|
| 171 |
+
"single_transformer_blocks.23.attn.to_out"
|
| 172 |
+
],
|
| 173 |
+
"adaln_modules": [
|
| 174 |
+
"double_stream_modulation_img.linear",
|
| 175 |
+
"double_stream_modulation_txt.linear",
|
| 176 |
+
"single_stream_modulation.linear"
|
| 177 |
+
],
|
| 178 |
+
"skipped_modules": [
|
| 179 |
+
"time_guidance_embed.timestep_embedder.linear_1",
|
| 180 |
+
"time_guidance_embed.timestep_embedder.linear_2",
|
| 181 |
+
"x_embedder",
|
| 182 |
+
"context_embedder",
|
| 183 |
+
"norm_out.linear",
|
| 184 |
+
"proj_out"
|
| 185 |
+
],
|
| 186 |
+
"quantization_device": "cuda",
|
| 187 |
+
"weight_quantization_backend": "triton_cuda",
|
| 188 |
+
"quantization_staging_mode": "component",
|
| 189 |
+
"orbitquant_seconds": 7.03080634213984,
|
| 190 |
+
"adaln_seconds": 1.2956953924149275,
|
| 191 |
+
"device_transfer_seconds": 3.2943784836679697,
|
| 192 |
+
"serialized_tensor_bytes": 4704718848,
|
| 193 |
+
"orbitquant_linear_count": 144,
|
| 194 |
+
"rtn_int4_linear_count": 3,
|
| 195 |
+
"cuda": {
|
| 196 |
+
"allocated_bytes": 4714399232,
|
| 197 |
+
"reserved_bytes": 18639486976,
|
| 198 |
+
"peak_allocated_bytes": 18619604992,
|
| 199 |
+
"peak_reserved_bytes": 18639486976
|
| 200 |
+
}
|
| 201 |
+
},
|
| 202 |
+
"text_encoder": {
|
| 203 |
+
"elapsed_seconds": 5.602843426167965,
|
| 204 |
+
"quantized_module_count": 252,
|
| 205 |
+
"adaln_module_count": 0,
|
| 206 |
+
"skipped_module_count": 1,
|
| 207 |
+
"quantized_modules": [
|
| 208 |
+
"model.layers.0.self_attn.q_proj",
|
| 209 |
+
"model.layers.0.self_attn.k_proj",
|
| 210 |
+
"model.layers.0.self_attn.v_proj",
|
| 211 |
+
"model.layers.0.self_attn.o_proj",
|
| 212 |
+
"model.layers.0.mlp.gate_proj",
|
| 213 |
+
"model.layers.0.mlp.up_proj",
|
| 214 |
+
"model.layers.0.mlp.down_proj",
|
| 215 |
+
"model.layers.1.self_attn.q_proj",
|
| 216 |
+
"model.layers.1.self_attn.k_proj",
|
| 217 |
+
"model.layers.1.self_attn.v_proj",
|
| 218 |
+
"model.layers.1.self_attn.o_proj",
|
| 219 |
+
"model.layers.1.mlp.gate_proj",
|
| 220 |
+
"model.layers.1.mlp.up_proj",
|
| 221 |
+
"model.layers.1.mlp.down_proj",
|
| 222 |
+
"model.layers.2.self_attn.q_proj",
|
| 223 |
+
"model.layers.2.self_attn.k_proj",
|
| 224 |
+
"model.layers.2.self_attn.v_proj",
|
| 225 |
+
"model.layers.2.self_attn.o_proj",
|
| 226 |
+
"model.layers.2.mlp.gate_proj",
|
| 227 |
+
"model.layers.2.mlp.up_proj",
|
| 228 |
+
"model.layers.2.mlp.down_proj",
|
| 229 |
+
"model.layers.3.self_attn.q_proj",
|
| 230 |
+
"model.layers.3.self_attn.k_proj",
|
| 231 |
+
"model.layers.3.self_attn.v_proj",
|
| 232 |
+
"model.layers.3.self_attn.o_proj",
|
| 233 |
+
"model.layers.3.mlp.gate_proj",
|
| 234 |
+
"model.layers.3.mlp.up_proj",
|
| 235 |
+
"model.layers.3.mlp.down_proj",
|
| 236 |
+
"model.layers.4.self_attn.q_proj",
|
| 237 |
+
"model.layers.4.self_attn.k_proj",
|
| 238 |
+
"model.layers.4.self_attn.v_proj",
|
| 239 |
+
"model.layers.4.self_attn.o_proj",
|
| 240 |
+
"model.layers.4.mlp.gate_proj",
|
| 241 |
+
"model.layers.4.mlp.up_proj",
|
| 242 |
+
"model.layers.4.mlp.down_proj",
|
| 243 |
+
"model.layers.5.self_attn.q_proj",
|
| 244 |
+
"model.layers.5.self_attn.k_proj",
|
| 245 |
+
"model.layers.5.self_attn.v_proj",
|
| 246 |
+
"model.layers.5.self_attn.o_proj",
|
| 247 |
+
"model.layers.5.mlp.gate_proj",
|
| 248 |
+
"model.layers.5.mlp.up_proj",
|
| 249 |
+
"model.layers.5.mlp.down_proj",
|
| 250 |
+
"model.layers.6.self_attn.q_proj",
|
| 251 |
+
"model.layers.6.self_attn.k_proj",
|
| 252 |
+
"model.layers.6.self_attn.v_proj",
|
| 253 |
+
"model.layers.6.self_attn.o_proj",
|
| 254 |
+
"model.layers.6.mlp.gate_proj",
|
| 255 |
+
"model.layers.6.mlp.up_proj",
|
| 256 |
+
"model.layers.6.mlp.down_proj",
|
| 257 |
+
"model.layers.7.self_attn.q_proj",
|
| 258 |
+
"model.layers.7.self_attn.k_proj",
|
| 259 |
+
"model.layers.7.self_attn.v_proj",
|
| 260 |
+
"model.layers.7.self_attn.o_proj",
|
| 261 |
+
"model.layers.7.mlp.gate_proj",
|
| 262 |
+
"model.layers.7.mlp.up_proj",
|
| 263 |
+
"model.layers.7.mlp.down_proj",
|
| 264 |
+
"model.layers.8.self_attn.q_proj",
|
| 265 |
+
"model.layers.8.self_attn.k_proj",
|
| 266 |
+
"model.layers.8.self_attn.v_proj",
|
| 267 |
+
"model.layers.8.self_attn.o_proj",
|
| 268 |
+
"model.layers.8.mlp.gate_proj",
|
| 269 |
+
"model.layers.8.mlp.up_proj",
|
| 270 |
+
"model.layers.8.mlp.down_proj",
|
| 271 |
+
"model.layers.9.self_attn.q_proj",
|
| 272 |
+
"model.layers.9.self_attn.k_proj",
|
| 273 |
+
"model.layers.9.self_attn.v_proj",
|
| 274 |
+
"model.layers.9.self_attn.o_proj",
|
| 275 |
+
"model.layers.9.mlp.gate_proj",
|
| 276 |
+
"model.layers.9.mlp.up_proj",
|
| 277 |
+
"model.layers.9.mlp.down_proj",
|
| 278 |
+
"model.layers.10.self_attn.q_proj",
|
| 279 |
+
"model.layers.10.self_attn.k_proj",
|
| 280 |
+
"model.layers.10.self_attn.v_proj",
|
| 281 |
+
"model.layers.10.self_attn.o_proj",
|
| 282 |
+
"model.layers.10.mlp.gate_proj",
|
| 283 |
+
"model.layers.10.mlp.up_proj",
|
| 284 |
+
"model.layers.10.mlp.down_proj",
|
| 285 |
+
"model.layers.11.self_attn.q_proj",
|
| 286 |
+
"model.layers.11.self_attn.k_proj",
|
| 287 |
+
"model.layers.11.self_attn.v_proj",
|
| 288 |
+
"model.layers.11.self_attn.o_proj",
|
| 289 |
+
"model.layers.11.mlp.gate_proj",
|
| 290 |
+
"model.layers.11.mlp.up_proj",
|
| 291 |
+
"model.layers.11.mlp.down_proj",
|
| 292 |
+
"model.layers.12.self_attn.q_proj",
|
| 293 |
+
"model.layers.12.self_attn.k_proj",
|
| 294 |
+
"model.layers.12.self_attn.v_proj",
|
| 295 |
+
"model.layers.12.self_attn.o_proj",
|
| 296 |
+
"model.layers.12.mlp.gate_proj",
|
| 297 |
+
"model.layers.12.mlp.up_proj",
|
| 298 |
+
"model.layers.12.mlp.down_proj",
|
| 299 |
+
"model.layers.13.self_attn.q_proj",
|
| 300 |
+
"model.layers.13.self_attn.k_proj",
|
| 301 |
+
"model.layers.13.self_attn.v_proj",
|
| 302 |
+
"model.layers.13.self_attn.o_proj",
|
| 303 |
+
"model.layers.13.mlp.gate_proj",
|
| 304 |
+
"model.layers.13.mlp.up_proj",
|
| 305 |
+
"model.layers.13.mlp.down_proj",
|
| 306 |
+
"model.layers.14.self_attn.q_proj",
|
| 307 |
+
"model.layers.14.self_attn.k_proj",
|
| 308 |
+
"model.layers.14.self_attn.v_proj",
|
| 309 |
+
"model.layers.14.self_attn.o_proj",
|
| 310 |
+
"model.layers.14.mlp.gate_proj",
|
| 311 |
+
"model.layers.14.mlp.up_proj",
|
| 312 |
+
"model.layers.14.mlp.down_proj",
|
| 313 |
+
"model.layers.15.self_attn.q_proj",
|
| 314 |
+
"model.layers.15.self_attn.k_proj",
|
| 315 |
+
"model.layers.15.self_attn.v_proj",
|
| 316 |
+
"model.layers.15.self_attn.o_proj",
|
| 317 |
+
"model.layers.15.mlp.gate_proj",
|
| 318 |
+
"model.layers.15.mlp.up_proj",
|
| 319 |
+
"model.layers.15.mlp.down_proj",
|
| 320 |
+
"model.layers.16.self_attn.q_proj",
|
| 321 |
+
"model.layers.16.self_attn.k_proj",
|
| 322 |
+
"model.layers.16.self_attn.v_proj",
|
| 323 |
+
"model.layers.16.self_attn.o_proj",
|
| 324 |
+
"model.layers.16.mlp.gate_proj",
|
| 325 |
+
"model.layers.16.mlp.up_proj",
|
| 326 |
+
"model.layers.16.mlp.down_proj",
|
| 327 |
+
"model.layers.17.self_attn.q_proj",
|
| 328 |
+
"model.layers.17.self_attn.k_proj",
|
| 329 |
+
"model.layers.17.self_attn.v_proj",
|
| 330 |
+
"model.layers.17.self_attn.o_proj",
|
| 331 |
+
"model.layers.17.mlp.gate_proj",
|
| 332 |
+
"model.layers.17.mlp.up_proj",
|
| 333 |
+
"model.layers.17.mlp.down_proj",
|
| 334 |
+
"model.layers.18.self_attn.q_proj",
|
| 335 |
+
"model.layers.18.self_attn.k_proj",
|
| 336 |
+
"model.layers.18.self_attn.v_proj",
|
| 337 |
+
"model.layers.18.self_attn.o_proj",
|
| 338 |
+
"model.layers.18.mlp.gate_proj",
|
| 339 |
+
"model.layers.18.mlp.up_proj",
|
| 340 |
+
"model.layers.18.mlp.down_proj",
|
| 341 |
+
"model.layers.19.self_attn.q_proj",
|
| 342 |
+
"model.layers.19.self_attn.k_proj",
|
| 343 |
+
"model.layers.19.self_attn.v_proj",
|
| 344 |
+
"model.layers.19.self_attn.o_proj",
|
| 345 |
+
"model.layers.19.mlp.gate_proj",
|
| 346 |
+
"model.layers.19.mlp.up_proj",
|
| 347 |
+
"model.layers.19.mlp.down_proj",
|
| 348 |
+
"model.layers.20.self_attn.q_proj",
|
| 349 |
+
"model.layers.20.self_attn.k_proj",
|
| 350 |
+
"model.layers.20.self_attn.v_proj",
|
| 351 |
+
"model.layers.20.self_attn.o_proj",
|
| 352 |
+
"model.layers.20.mlp.gate_proj",
|
| 353 |
+
"model.layers.20.mlp.up_proj",
|
| 354 |
+
"model.layers.20.mlp.down_proj",
|
| 355 |
+
"model.layers.21.self_attn.q_proj",
|
| 356 |
+
"model.layers.21.self_attn.k_proj",
|
| 357 |
+
"model.layers.21.self_attn.v_proj",
|
| 358 |
+
"model.layers.21.self_attn.o_proj",
|
| 359 |
+
"model.layers.21.mlp.gate_proj",
|
| 360 |
+
"model.layers.21.mlp.up_proj",
|
| 361 |
+
"model.layers.21.mlp.down_proj",
|
| 362 |
+
"model.layers.22.self_attn.q_proj",
|
| 363 |
+
"model.layers.22.self_attn.k_proj",
|
| 364 |
+
"model.layers.22.self_attn.v_proj",
|
| 365 |
+
"model.layers.22.self_attn.o_proj",
|
| 366 |
+
"model.layers.22.mlp.gate_proj",
|
| 367 |
+
"model.layers.22.mlp.up_proj",
|
| 368 |
+
"model.layers.22.mlp.down_proj",
|
| 369 |
+
"model.layers.23.self_attn.q_proj",
|
| 370 |
+
"model.layers.23.self_attn.k_proj",
|
| 371 |
+
"model.layers.23.self_attn.v_proj",
|
| 372 |
+
"model.layers.23.self_attn.o_proj",
|
| 373 |
+
"model.layers.23.mlp.gate_proj",
|
| 374 |
+
"model.layers.23.mlp.up_proj",
|
| 375 |
+
"model.layers.23.mlp.down_proj",
|
| 376 |
+
"model.layers.24.self_attn.q_proj",
|
| 377 |
+
"model.layers.24.self_attn.k_proj",
|
| 378 |
+
"model.layers.24.self_attn.v_proj",
|
| 379 |
+
"model.layers.24.self_attn.o_proj",
|
| 380 |
+
"model.layers.24.mlp.gate_proj",
|
| 381 |
+
"model.layers.24.mlp.up_proj",
|
| 382 |
+
"model.layers.24.mlp.down_proj",
|
| 383 |
+
"model.layers.25.self_attn.q_proj",
|
| 384 |
+
"model.layers.25.self_attn.k_proj",
|
| 385 |
+
"model.layers.25.self_attn.v_proj",
|
| 386 |
+
"model.layers.25.self_attn.o_proj",
|
| 387 |
+
"model.layers.25.mlp.gate_proj",
|
| 388 |
+
"model.layers.25.mlp.up_proj",
|
| 389 |
+
"model.layers.25.mlp.down_proj",
|
| 390 |
+
"model.layers.26.self_attn.q_proj",
|
| 391 |
+
"model.layers.26.self_attn.k_proj",
|
| 392 |
+
"model.layers.26.self_attn.v_proj",
|
| 393 |
+
"model.layers.26.self_attn.o_proj",
|
| 394 |
+
"model.layers.26.mlp.gate_proj",
|
| 395 |
+
"model.layers.26.mlp.up_proj",
|
| 396 |
+
"model.layers.26.mlp.down_proj",
|
| 397 |
+
"model.layers.27.self_attn.q_proj",
|
| 398 |
+
"model.layers.27.self_attn.k_proj",
|
| 399 |
+
"model.layers.27.self_attn.v_proj",
|
| 400 |
+
"model.layers.27.self_attn.o_proj",
|
| 401 |
+
"model.layers.27.mlp.gate_proj",
|
| 402 |
+
"model.layers.27.mlp.up_proj",
|
| 403 |
+
"model.layers.27.mlp.down_proj",
|
| 404 |
+
"model.layers.28.self_attn.q_proj",
|
| 405 |
+
"model.layers.28.self_attn.k_proj",
|
| 406 |
+
"model.layers.28.self_attn.v_proj",
|
| 407 |
+
"model.layers.28.self_attn.o_proj",
|
| 408 |
+
"model.layers.28.mlp.gate_proj",
|
| 409 |
+
"model.layers.28.mlp.up_proj",
|
| 410 |
+
"model.layers.28.mlp.down_proj",
|
| 411 |
+
"model.layers.29.self_attn.q_proj",
|
| 412 |
+
"model.layers.29.self_attn.k_proj",
|
| 413 |
+
"model.layers.29.self_attn.v_proj",
|
| 414 |
+
"model.layers.29.self_attn.o_proj",
|
| 415 |
+
"model.layers.29.mlp.gate_proj",
|
| 416 |
+
"model.layers.29.mlp.up_proj",
|
| 417 |
+
"model.layers.29.mlp.down_proj",
|
| 418 |
+
"model.layers.30.self_attn.q_proj",
|
| 419 |
+
"model.layers.30.self_attn.k_proj",
|
| 420 |
+
"model.layers.30.self_attn.v_proj",
|
| 421 |
+
"model.layers.30.self_attn.o_proj",
|
| 422 |
+
"model.layers.30.mlp.gate_proj",
|
| 423 |
+
"model.layers.30.mlp.up_proj",
|
| 424 |
+
"model.layers.30.mlp.down_proj",
|
| 425 |
+
"model.layers.31.self_attn.q_proj",
|
| 426 |
+
"model.layers.31.self_attn.k_proj",
|
| 427 |
+
"model.layers.31.self_attn.v_proj",
|
| 428 |
+
"model.layers.31.self_attn.o_proj",
|
| 429 |
+
"model.layers.31.mlp.gate_proj",
|
| 430 |
+
"model.layers.31.mlp.up_proj",
|
| 431 |
+
"model.layers.31.mlp.down_proj",
|
| 432 |
+
"model.layers.32.self_attn.q_proj",
|
| 433 |
+
"model.layers.32.self_attn.k_proj",
|
| 434 |
+
"model.layers.32.self_attn.v_proj",
|
| 435 |
+
"model.layers.32.self_attn.o_proj",
|
| 436 |
+
"model.layers.32.mlp.gate_proj",
|
| 437 |
+
"model.layers.32.mlp.up_proj",
|
| 438 |
+
"model.layers.32.mlp.down_proj",
|
| 439 |
+
"model.layers.33.self_attn.q_proj",
|
| 440 |
+
"model.layers.33.self_attn.k_proj",
|
| 441 |
+
"model.layers.33.self_attn.v_proj",
|
| 442 |
+
"model.layers.33.self_attn.o_proj",
|
| 443 |
+
"model.layers.33.mlp.gate_proj",
|
| 444 |
+
"model.layers.33.mlp.up_proj",
|
| 445 |
+
"model.layers.33.mlp.down_proj",
|
| 446 |
+
"model.layers.34.self_attn.q_proj",
|
| 447 |
+
"model.layers.34.self_attn.k_proj",
|
| 448 |
+
"model.layers.34.self_attn.v_proj",
|
| 449 |
+
"model.layers.34.self_attn.o_proj",
|
| 450 |
+
"model.layers.34.mlp.gate_proj",
|
| 451 |
+
"model.layers.34.mlp.up_proj",
|
| 452 |
+
"model.layers.34.mlp.down_proj",
|
| 453 |
+
"model.layers.35.self_attn.q_proj",
|
| 454 |
+
"model.layers.35.self_attn.k_proj",
|
| 455 |
+
"model.layers.35.self_attn.v_proj",
|
| 456 |
+
"model.layers.35.self_attn.o_proj",
|
| 457 |
+
"model.layers.35.mlp.gate_proj",
|
| 458 |
+
"model.layers.35.mlp.up_proj",
|
| 459 |
+
"model.layers.35.mlp.down_proj"
|
| 460 |
+
],
|
| 461 |
+
"adaln_modules": [],
|
| 462 |
+
"skipped_modules": [
|
| 463 |
+
"lm_head"
|
| 464 |
+
],
|
| 465 |
+
"quantization_device": "cuda",
|
| 466 |
+
"weight_quantization_backend": "triton_cuda",
|
| 467 |
+
"quantization_staging_mode": "component",
|
| 468 |
+
"orbitquant_seconds": 3.190608648583293,
|
| 469 |
+
"adaln_seconds": 0.0,
|
| 470 |
+
"device_transfer_seconds": 2.393733749166131,
|
| 471 |
+
"serialized_tensor_bytes": 5965621248,
|
| 472 |
+
"orbitquant_linear_count": 252,
|
| 473 |
+
"rtn_int4_linear_count": 0,
|
| 474 |
+
"cuda": {
|
| 475 |
+
"allocated_bytes": 5978268160,
|
| 476 |
+
"reserved_bytes": 16626221056,
|
| 477 |
+
"peak_allocated_bytes": 16545587712,
|
| 478 |
+
"peak_reserved_bytes": 16626221056
|
| 479 |
+
}
|
| 480 |
+
}
|
| 481 |
+
},
|
| 482 |
+
"build": {
|
| 483 |
+
"source_load_seconds": 17.03382627852261,
|
| 484 |
+
"save_seconds": 24.1477939337492,
|
| 485 |
+
"total_seconds": 67.16633238084614,
|
| 486 |
+
"process_rss_bytes": 12405612544,
|
| 487 |
+
"process_peak_rss_bytes": 12405612544,
|
| 488 |
+
"torch_version": "2.9.1+cu128",
|
| 489 |
+
"cuda_version": "12.8",
|
| 490 |
+
"gpu": "NVIDIA A40"
|
| 491 |
+
},
|
| 492 |
+
"files": {
|
| 493 |
+
"prompts.json": 10279,
|
| 494 |
+
"LICENSE.md": 18158,
|
| 495 |
+
"model_index.json": 493,
|
| 496 |
+
"transformer/diffusion_pytorch_model.safetensors": 4704766296,
|
| 497 |
+
"transformer/config.json": 1487,
|
| 498 |
+
"scheduler/scheduler_config.json": 481,
|
| 499 |
+
"tokenizer/tokenizer.json": 11422650,
|
| 500 |
+
"tokenizer/tokenizer_config.json": 377,
|
| 501 |
+
"tokenizer/chat_template.jinja": 4168,
|
| 502 |
+
"text_encoder/model.safetensors.index.json": 57899,
|
| 503 |
+
"text_encoder/model-00002-of-00002.safetensors": 965657624,
|
| 504 |
+
"text_encoder/model-00001-of-00002.safetensors": 5000040328,
|
| 505 |
+
"text_encoder/generation_config.json": 142,
|
| 506 |
+
"text_encoder/config.json": 2474,
|
| 507 |
+
"vae/diffusion_pytorch_model.safetensors": 168120878,
|
| 508 |
+
"vae/config.json": 940
|
| 509 |
+
},
|
| 510 |
+
"total_file_bytes": 10850104674,
|
| 511 |
+
"minimum_runtime_version": "orbitquant>=0.2.2",
|
| 512 |
+
"benchmark_summary": "benchmark/summary.json",
|
| 513 |
+
"comparison_asset": "assets/image_generation_comparison_matrix.webp"
|
| 514 |
+
}
|
prompts.json
ADDED
|
@@ -0,0 +1,104 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"prompt_pack": "flux2_klein_9b_quantizer_stress_v1",
|
| 3 |
+
"prompts": [
|
| 4 |
+
{
|
| 5 |
+
"id": "micro-mechanical-cathedral",
|
| 6 |
+
"title": "01 Micro-mechanical cathedral reliquary",
|
| 7 |
+
"category": "extreme_micro_detail",
|
| 8 |
+
"prompt": "A museum-grade large-format macro photograph of a hand-sized brass reliquary built as a complete Gothic cathedral, its rose windows made from translucent enamel and its nave filled with hundreds of functional watch gears, jeweled bearings, hair-thin chains, engraved astronomical tables, tiny saints no taller than a grain of rice, minute oxidized solder seams, fingerprints in old wax, and dust caught between moving teeth; one door is open to reveal a second clockwork chapel inside, physically coherent mechanisms, razor-sharp focus stacking, black velvet background, Rembrandt light, controlled specular highlights, archival object photography, no duplicated gears or melted ornament",
|
| 9 |
+
"required_text": []
|
| 10 |
+
},
|
| 11 |
+
{
|
| 12 |
+
"id": "nine-character-opera",
|
| 13 |
+
"title": "02 Nine-character flooded opera composition",
|
| 14 |
+
"category": "counting_spatial_character_consistency",
|
| 15 |
+
"prompt": "Exactly nine distinct performers rehearse in a flooded baroque opera house, each standing on a separate circular platform and no additional people anywhere: three crimson masked dancers on the left, three ivory masked dancers in the center, and three cobalt masked dancers on the right; the center performer holds a silver astrolabe above her head, the far-left performer kneels beside a mechanical swan, and the far-right performer releases a paper bird; balconies, chandeliers, costumes, and all nine bodies reflect correctly in dark water, floating candles pass between platforms, stage fog has believable depth, intricate hands and faces, symmetrical yet cinematic wide composition, deep focus, no crowd",
|
| 16 |
+
"required_text": []
|
| 17 |
+
},
|
| 18 |
+
{
|
| 19 |
+
"id": "vertical-city-cutaway",
|
| 20 |
+
"title": "03 Nested vertical-city architectural cutaway",
|
| 21 |
+
"category": "complex_nested_composition",
|
| 22 |
+
"prompt": "An impossibly detailed isometric architectural cutaway of a vertical city at dusk: a botanical laboratory occupies the glass roof directly above an elevated silver train; beneath the train is a circular library, beneath the library is a flooded market, and beneath the market an orange submarine docks in a cavern; a red fox stands inside the rooftop lab under a moon lamp, a violinist waits on the train platform, miniature chefs work in the market, and a yellow airship passes behind the entire structure; hundreds of coherent rooms, stairs, pipes, plants, cables, windows and tiny inhabitants, every floor relationship readable, photoreal materials mixed with precise architectural-section drawing, clean perspective, soft volumetric twilight, no floating disconnected rooms",
|
| 23 |
+
"required_text": []
|
| 24 |
+
},
|
| 25 |
+
{
|
| 26 |
+
"id": "fictional-auteur-collage",
|
| 27 |
+
"title": "04 Fictional auteur mixed-media collage",
|
| 28 |
+
"category": "original_authorial_style",
|
| 29 |
+
"prompt": "A monumental mixed-media artwork attributed to a fictional avant-garde artist named Irena Volskaya, whose invented signature language combines surgical botanical diagrams, translucent mineral pigments, hand-stitched copper wire, torn meteorological maps, lacquered black voids, and repeated white ladder motifs; the composition depicts a city remembering its own demolition, with recognizable buildings dissolving into anatomical flowers and weather systems, a tiny red observer repeated exactly five times, deliberate asymmetric balance, visible canvas weave, cracked gesso, ink bleed, embossed paper edges and gallery lighting; original auteur work, emotionally severe but visually controlled, neither generic surrealism nor a copy of any real artist",
|
| 30 |
+
"required_text": []
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"id": "abstract-material-triptych",
|
| 34 |
+
"title": "05 Materially complex abstract triptych",
|
| 35 |
+
"category": "abstract_art_geometry_materials",
|
| 36 |
+
"prompt": "A three-panel abstract museum installation titled Silent Tectonics: the left panel contains seven ultramarine arcs interrupted by one vermilion square, the center panel contains a suspended translucent amber spiral crossing a field of graphite dust, and the right panel contains twelve porcelain-white vertical cuts over oxidized green copper; connect all three panels with one continuous hair-thin gold line that changes depth without breaking; encaustic wax, sanded aluminum, soot, silk fibers, iridescent resin and hand-burnished leaf must remain materially distinct, subtle relief shadows, restrained contemporary gallery, mathematically intentional negative space, high-resolution art documentation, no symbols, faces or representational objects",
|
| 37 |
+
"required_text": []
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"id": "english-micro-typography",
|
| 41 |
+
"title": "06 English technical poster with micro typography",
|
| 42 |
+
"category": "latin_large_and_small_typography",
|
| 43 |
+
"prompt": "A pristine Swiss technical exhibition poster photographed flat under museum light, with the exact large headline \"ORBITAL MEMORY\", exact subtitle \"DATA WITHOUT CALIBRATION\", and a fine-print specification table containing exactly these four readable lines: \"ROTATION: RPBH\", \"WEIGHTS: FOUR BIT\", \"ACTIVATIONS: FOUR BIT\", and \"REVISION: 2049-A\"; also include tiny axis labels \"TIME\", \"SIGNAL\", \"NOISE\", and \"PHASE\" around a precise orbital diagram; strict modular grid, black, white, red and electric blue screenprint, registration marks, embossed cotton paper, sharp kerning at every scale, no additional words, no misspellings",
|
| 44 |
+
"required_text": [
|
| 45 |
+
"ORBITAL MEMORY",
|
| 46 |
+
"DATA WITHOUT CALIBRATION",
|
| 47 |
+
"ROTATION: RPBH",
|
| 48 |
+
"WEIGHTS: FOUR BIT",
|
| 49 |
+
"ACTIVATIONS: FOUR BIT",
|
| 50 |
+
"REVISION: 2049-A"
|
| 51 |
+
]
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"id": "russian-editorial-typography",
|
| 55 |
+
"title": "07 Russian editorial page with fine print",
|
| 56 |
+
"category": "cyrillic_large_and_small_typography",
|
| 57 |
+
"prompt": "A richly layered Russian Constructivist science journal cover printed on aged cream stock, with the exact headline \"КВАНТОВАЯ ОРБИТА\", exact subtitle \"МОСКВА 2049\", a red circular stamp reading \"АРХИВ\", and a small but legible contents column containing exactly \"МАТЕРИЯ И СВЕТ\", \"ПАМЯТЬ МАШИН\", \"КАРТА ВРЕМЕНИ\", and \"ВЫПУСК 07\"; diagonal black and scarlet geometry surrounds a cosmonaut repairing a transparent orbital computer, with fine technical callouts, halftone portraits, folded corners, ink trapping, paper fibers and slight plate misregistration; every Cyrillic letter crisp and correctly ordered, no invented alphabet and no extra text",
|
| 58 |
+
"required_text": [
|
| 59 |
+
"КВАНТОВАЯ ОРБИТА",
|
| 60 |
+
"МОСКВА 2049",
|
| 61 |
+
"АРХИВ",
|
| 62 |
+
"МАТЕРИЯ И СВЕТ",
|
| 63 |
+
"ПАМЯТЬ МАШИН",
|
| 64 |
+
"КАРТА ВРЕМЕНИ",
|
| 65 |
+
"ВЫПУСК 07"
|
| 66 |
+
]
|
| 67 |
+
},
|
| 68 |
+
{
|
| 69 |
+
"id": "japanese-magazine-typography",
|
| 70 |
+
"title": "08 Japanese magazine cover and small labels",
|
| 71 |
+
"category": "japanese_typography_mixed_style",
|
| 72 |
+
"prompt": "An elaborate Japanese architecture magazine cover combining Edo woodblock printing, precision photography and a futuristic Tokyo transit atlas, with the exact vertical title \"量子の軌道\", exact subtitle \"東京の未来\", and four small readable cover lines \"光と記憶\", \"都市の断面\", \"機械の庭\", and \"特集 2049\"; giant indigo waves curl around glass towers while red-crowned cranes cross a gold moon, miniature trains, signs, pedestrians and rooftop gardens fill the lower city, visible washi fibers, layered spot colors, foil accents, disciplined editorial hierarchy, precise Japanese glyphs, no pseudo-characters and no additional text",
|
| 73 |
+
"required_text": [
|
| 74 |
+
"量子の軌道",
|
| 75 |
+
"東京の未来",
|
| 76 |
+
"光と記憶",
|
| 77 |
+
"都市の断面",
|
| 78 |
+
"機械の庭",
|
| 79 |
+
"特集 2049"
|
| 80 |
+
]
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"id": "chinese-reflective-storefront",
|
| 84 |
+
"title": "09 Chinese storefront typography and reflections",
|
| 85 |
+
"category": "chinese_typography_reflection_occlusion",
|
| 86 |
+
"prompt": "A luxurious Chinese retro-futurist apothecary window on a rain-soaked Shanghai street, with the exact gold sign \"量子轨道\", exact red subtitle \"未来之城\", and tiny readable drawer labels \"月光\", \"记忆\", \"时间\", and \"星尘\"; a curved chrome medical robot is partly occluded by peonies, blue-and-white porcelain and hanging silk, while the main sign, passing bicycles, neon street and drawer labels reflect coherently across its body and three layers of glass; hundreds of medicinal jars and brass mechanisms, wet pavement, cinematic rain, subtle interior figures, meticulous product photography, all Chinese characters correctly formed, no extra writing",
|
| 87 |
+
"required_text": [
|
| 88 |
+
"量子轨道",
|
| 89 |
+
"未来之城",
|
| 90 |
+
"月光",
|
| 91 |
+
"记忆",
|
| 92 |
+
"时间",
|
| 93 |
+
"星尘"
|
| 94 |
+
]
|
| 95 |
+
},
|
| 96 |
+
{
|
| 97 |
+
"id": "orbital-banquet-panorama",
|
| 98 |
+
"title": "10 Orbital banquet panoramic master composition",
|
| 99 |
+
"category": "maximum_complexity_panorama",
|
| 100 |
+
"prompt": "A sweeping anamorphic panorama inside a rotating glass orbital conservatory during an impossible diplomatic banquet: in the immediate foreground a chef plates translucent dumplings beside a cracked mirror and silver cutlery; behind him a masked string quartet performs on a bridge crossing a koi pond; to the left, botanists release luminous moths among giant orchids; to the right, two diplomats exchange a miniature mechanical planet while a child watches through an aquarium; farther back, hundreds of guests circulate through terraced gardens, service robots carry lanterns, and Earth rises behind the curved windows; every foreground object must occlude the correct background layer, mirror, water, chrome and glass reflections must agree, coherent faces and hands, deep focus, minute table settings, embroidered fabrics, condensation, cables, leaves and distant spacecraft, painterly color design with documentary realism and no empty generic regions",
|
| 101 |
+
"required_text": []
|
| 102 |
+
}
|
| 103 |
+
]
|
| 104 |
+
}
|
scheduler/scheduler_config.json
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_class_name": "FlowMatchEulerDiscreteScheduler",
|
| 3 |
+
"_diffusers_version": "0.39.0",
|
| 4 |
+
"base_image_seq_len": 256,
|
| 5 |
+
"base_shift": 0.5,
|
| 6 |
+
"invert_sigmas": false,
|
| 7 |
+
"max_image_seq_len": 4096,
|
| 8 |
+
"max_shift": 1.15,
|
| 9 |
+
"num_train_timesteps": 1000,
|
| 10 |
+
"shift": 3.0,
|
| 11 |
+
"shift_terminal": null,
|
| 12 |
+
"stochastic_sampling": false,
|
| 13 |
+
"time_shift_type": "exponential",
|
| 14 |
+
"use_beta_sigmas": false,
|
| 15 |
+
"use_dynamic_shifting": true,
|
| 16 |
+
"use_exponential_sigmas": false,
|
| 17 |
+
"use_karras_sigmas": false
|
| 18 |
+
}
|
text_encoder/config.json
ADDED
|
@@ -0,0 +1,100 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Qwen3ForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"attention_bias": false,
|
| 6 |
+
"attention_dropout": 0.0,
|
| 7 |
+
"bos_token_id": 151643,
|
| 8 |
+
"dtype": "bfloat16",
|
| 9 |
+
"eos_token_id": 151645,
|
| 10 |
+
"head_dim": 128,
|
| 11 |
+
"hidden_act": "silu",
|
| 12 |
+
"hidden_size": 4096,
|
| 13 |
+
"initializer_range": 0.02,
|
| 14 |
+
"intermediate_size": 12288,
|
| 15 |
+
"layer_types": [
|
| 16 |
+
"full_attention",
|
| 17 |
+
"full_attention",
|
| 18 |
+
"full_attention",
|
| 19 |
+
"full_attention",
|
| 20 |
+
"full_attention",
|
| 21 |
+
"full_attention",
|
| 22 |
+
"full_attention",
|
| 23 |
+
"full_attention",
|
| 24 |
+
"full_attention",
|
| 25 |
+
"full_attention",
|
| 26 |
+
"full_attention",
|
| 27 |
+
"full_attention",
|
| 28 |
+
"full_attention",
|
| 29 |
+
"full_attention",
|
| 30 |
+
"full_attention",
|
| 31 |
+
"full_attention",
|
| 32 |
+
"full_attention",
|
| 33 |
+
"full_attention",
|
| 34 |
+
"full_attention",
|
| 35 |
+
"full_attention",
|
| 36 |
+
"full_attention",
|
| 37 |
+
"full_attention",
|
| 38 |
+
"full_attention",
|
| 39 |
+
"full_attention",
|
| 40 |
+
"full_attention",
|
| 41 |
+
"full_attention",
|
| 42 |
+
"full_attention",
|
| 43 |
+
"full_attention",
|
| 44 |
+
"full_attention",
|
| 45 |
+
"full_attention",
|
| 46 |
+
"full_attention",
|
| 47 |
+
"full_attention",
|
| 48 |
+
"full_attention",
|
| 49 |
+
"full_attention",
|
| 50 |
+
"full_attention",
|
| 51 |
+
"full_attention"
|
| 52 |
+
],
|
| 53 |
+
"max_position_embeddings": 40960,
|
| 54 |
+
"max_window_layers": 36,
|
| 55 |
+
"model_type": "qwen3",
|
| 56 |
+
"num_attention_heads": 32,
|
| 57 |
+
"num_hidden_layers": 36,
|
| 58 |
+
"num_key_value_heads": 8,
|
| 59 |
+
"pad_token_id": null,
|
| 60 |
+
"quantization_config": {
|
| 61 |
+
"activation_bits": 4,
|
| 62 |
+
"activation_eps": 1e-10,
|
| 63 |
+
"activation_kernel_backend": "triton_cuda",
|
| 64 |
+
"activation_norm_dtype": "float32",
|
| 65 |
+
"adaln_group_size": 64,
|
| 66 |
+
"adaln_policy": "int4_rtn",
|
| 67 |
+
"artifact_format_version": 1,
|
| 68 |
+
"block_size": "paper",
|
| 69 |
+
"codebook": "lloyd_max",
|
| 70 |
+
"codebook_dtype": "float32",
|
| 71 |
+
"codebook_version": 2,
|
| 72 |
+
"modules_dtype_dict": {},
|
| 73 |
+
"modules_to_convert": [],
|
| 74 |
+
"modules_to_not_convert": [],
|
| 75 |
+
"modules_to_use_adaln": [],
|
| 76 |
+
"packed_matmul_block_k": 128,
|
| 77 |
+
"packed_matmul_block_m": 64,
|
| 78 |
+
"packed_matmul_block_n": 64,
|
| 79 |
+
"packed_matmul_num_warps": 4,
|
| 80 |
+
"quant_method": "orbitquant",
|
| 81 |
+
"rotation": "rpbh",
|
| 82 |
+
"rotation_seed": 0,
|
| 83 |
+
"row_norm_dtype": "bfloat16",
|
| 84 |
+
"runtime_mode": "auto_fused",
|
| 85 |
+
"target_policy": "auto",
|
| 86 |
+
"weight_bits": 4,
|
| 87 |
+
"weight_pack_dtype": "uint8"
|
| 88 |
+
},
|
| 89 |
+
"rms_norm_eps": 1e-06,
|
| 90 |
+
"rope_parameters": {
|
| 91 |
+
"rope_theta": 1000000,
|
| 92 |
+
"rope_type": "default"
|
| 93 |
+
},
|
| 94 |
+
"sliding_window": null,
|
| 95 |
+
"tie_word_embeddings": false,
|
| 96 |
+
"transformers_version": "5.13.0",
|
| 97 |
+
"use_cache": true,
|
| 98 |
+
"use_sliding_window": false,
|
| 99 |
+
"vocab_size": 151936
|
| 100 |
+
}
|
text_encoder/generation_config.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_from_model_config": true,
|
| 3 |
+
"bos_token_id": 151643,
|
| 4 |
+
"eos_token_id": 151645,
|
| 5 |
+
"transformers_version": "5.13.0",
|
| 6 |
+
"use_cache": true
|
| 7 |
+
}
|
text_encoder/model-00001-of-00002.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:baff83ea9d23ff652ef1b0647e7dfd0af153797183a6edf9098740b168067253
|
| 3 |
+
size 5000040328
|
text_encoder/model-00002-of-00002.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:147e9d7ecd892c90d5fecedfa0a3e0706b5bd234b2d27a8071fcc1508c29b834
|
| 3 |
+
size 965657624
|
text_encoder/model.safetensors.index.json
ADDED
|
@@ -0,0 +1,659 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"metadata": {
|
| 3 |
+
"total_parameters": 1244967936,
|
| 4 |
+
"total_size": 5965621248
|
| 5 |
+
},
|
| 6 |
+
"weight_map": {
|
| 7 |
+
"lm_head.weight": "model-00001-of-00002.safetensors",
|
| 8 |
+
"model.embed_tokens.weight": "model-00001-of-00002.safetensors",
|
| 9 |
+
"model.layers.0.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 10 |
+
"model.layers.0.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 11 |
+
"model.layers.0.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 12 |
+
"model.layers.0.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 13 |
+
"model.layers.0.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 14 |
+
"model.layers.0.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 15 |
+
"model.layers.0.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 16 |
+
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 17 |
+
"model.layers.0.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 18 |
+
"model.layers.0.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 19 |
+
"model.layers.0.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 20 |
+
"model.layers.0.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 21 |
+
"model.layers.0.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 22 |
+
"model.layers.0.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 23 |
+
"model.layers.0.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 24 |
+
"model.layers.0.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 25 |
+
"model.layers.0.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 26 |
+
"model.layers.0.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 27 |
+
"model.layers.1.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 28 |
+
"model.layers.1.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 29 |
+
"model.layers.1.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 30 |
+
"model.layers.1.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 31 |
+
"model.layers.1.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 32 |
+
"model.layers.1.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 33 |
+
"model.layers.1.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 34 |
+
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 35 |
+
"model.layers.1.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 36 |
+
"model.layers.1.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 37 |
+
"model.layers.1.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 38 |
+
"model.layers.1.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 39 |
+
"model.layers.1.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 40 |
+
"model.layers.1.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 41 |
+
"model.layers.1.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 42 |
+
"model.layers.1.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 43 |
+
"model.layers.1.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 44 |
+
"model.layers.1.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 45 |
+
"model.layers.10.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 46 |
+
"model.layers.10.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 47 |
+
"model.layers.10.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 48 |
+
"model.layers.10.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 49 |
+
"model.layers.10.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 50 |
+
"model.layers.10.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 51 |
+
"model.layers.10.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 52 |
+
"model.layers.10.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 53 |
+
"model.layers.10.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 54 |
+
"model.layers.10.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 55 |
+
"model.layers.10.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 56 |
+
"model.layers.10.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 57 |
+
"model.layers.10.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 58 |
+
"model.layers.10.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 59 |
+
"model.layers.10.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 60 |
+
"model.layers.10.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 61 |
+
"model.layers.10.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 62 |
+
"model.layers.10.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 63 |
+
"model.layers.11.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 64 |
+
"model.layers.11.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 65 |
+
"model.layers.11.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 66 |
+
"model.layers.11.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 67 |
+
"model.layers.11.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 68 |
+
"model.layers.11.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 69 |
+
"model.layers.11.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 70 |
+
"model.layers.11.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 71 |
+
"model.layers.11.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 72 |
+
"model.layers.11.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 73 |
+
"model.layers.11.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 74 |
+
"model.layers.11.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 75 |
+
"model.layers.11.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 76 |
+
"model.layers.11.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 77 |
+
"model.layers.11.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 78 |
+
"model.layers.11.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 79 |
+
"model.layers.11.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 80 |
+
"model.layers.11.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 81 |
+
"model.layers.12.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 82 |
+
"model.layers.12.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 83 |
+
"model.layers.12.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 84 |
+
"model.layers.12.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 85 |
+
"model.layers.12.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 86 |
+
"model.layers.12.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 87 |
+
"model.layers.12.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 88 |
+
"model.layers.12.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 89 |
+
"model.layers.12.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 90 |
+
"model.layers.12.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 91 |
+
"model.layers.12.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 92 |
+
"model.layers.12.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 93 |
+
"model.layers.12.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 94 |
+
"model.layers.12.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 95 |
+
"model.layers.12.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 96 |
+
"model.layers.12.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 97 |
+
"model.layers.12.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 98 |
+
"model.layers.12.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 99 |
+
"model.layers.13.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 100 |
+
"model.layers.13.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 101 |
+
"model.layers.13.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 102 |
+
"model.layers.13.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 103 |
+
"model.layers.13.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 104 |
+
"model.layers.13.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 105 |
+
"model.layers.13.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 106 |
+
"model.layers.13.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 107 |
+
"model.layers.13.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 108 |
+
"model.layers.13.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 109 |
+
"model.layers.13.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 110 |
+
"model.layers.13.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 111 |
+
"model.layers.13.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 112 |
+
"model.layers.13.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 113 |
+
"model.layers.13.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 114 |
+
"model.layers.13.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 115 |
+
"model.layers.13.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 116 |
+
"model.layers.13.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 117 |
+
"model.layers.14.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 118 |
+
"model.layers.14.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 119 |
+
"model.layers.14.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 120 |
+
"model.layers.14.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 121 |
+
"model.layers.14.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 122 |
+
"model.layers.14.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 123 |
+
"model.layers.14.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 124 |
+
"model.layers.14.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 125 |
+
"model.layers.14.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 126 |
+
"model.layers.14.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 127 |
+
"model.layers.14.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 128 |
+
"model.layers.14.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 129 |
+
"model.layers.14.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 130 |
+
"model.layers.14.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 131 |
+
"model.layers.14.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 132 |
+
"model.layers.14.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 133 |
+
"model.layers.14.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 134 |
+
"model.layers.14.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 135 |
+
"model.layers.15.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 136 |
+
"model.layers.15.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 137 |
+
"model.layers.15.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 138 |
+
"model.layers.15.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 139 |
+
"model.layers.15.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 140 |
+
"model.layers.15.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 141 |
+
"model.layers.15.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 142 |
+
"model.layers.15.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 143 |
+
"model.layers.15.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 144 |
+
"model.layers.15.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 145 |
+
"model.layers.15.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 146 |
+
"model.layers.15.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 147 |
+
"model.layers.15.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 148 |
+
"model.layers.15.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 149 |
+
"model.layers.15.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 150 |
+
"model.layers.15.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 151 |
+
"model.layers.15.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 152 |
+
"model.layers.15.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 153 |
+
"model.layers.16.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 154 |
+
"model.layers.16.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 155 |
+
"model.layers.16.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 156 |
+
"model.layers.16.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 157 |
+
"model.layers.16.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 158 |
+
"model.layers.16.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 159 |
+
"model.layers.16.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 160 |
+
"model.layers.16.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 161 |
+
"model.layers.16.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 162 |
+
"model.layers.16.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 163 |
+
"model.layers.16.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 164 |
+
"model.layers.16.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 165 |
+
"model.layers.16.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 166 |
+
"model.layers.16.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 167 |
+
"model.layers.16.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 168 |
+
"model.layers.16.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 169 |
+
"model.layers.16.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 170 |
+
"model.layers.16.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 171 |
+
"model.layers.17.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 172 |
+
"model.layers.17.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 173 |
+
"model.layers.17.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 174 |
+
"model.layers.17.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 175 |
+
"model.layers.17.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 176 |
+
"model.layers.17.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 177 |
+
"model.layers.17.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 178 |
+
"model.layers.17.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 179 |
+
"model.layers.17.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 180 |
+
"model.layers.17.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 181 |
+
"model.layers.17.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 182 |
+
"model.layers.17.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 183 |
+
"model.layers.17.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 184 |
+
"model.layers.17.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 185 |
+
"model.layers.17.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 186 |
+
"model.layers.17.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 187 |
+
"model.layers.17.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 188 |
+
"model.layers.17.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 189 |
+
"model.layers.18.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 190 |
+
"model.layers.18.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 191 |
+
"model.layers.18.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 192 |
+
"model.layers.18.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 193 |
+
"model.layers.18.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 194 |
+
"model.layers.18.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 195 |
+
"model.layers.18.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 196 |
+
"model.layers.18.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 197 |
+
"model.layers.18.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 198 |
+
"model.layers.18.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 199 |
+
"model.layers.18.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 200 |
+
"model.layers.18.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 201 |
+
"model.layers.18.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 202 |
+
"model.layers.18.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 203 |
+
"model.layers.18.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 204 |
+
"model.layers.18.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 205 |
+
"model.layers.18.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 206 |
+
"model.layers.18.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 207 |
+
"model.layers.19.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 208 |
+
"model.layers.19.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 209 |
+
"model.layers.19.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 210 |
+
"model.layers.19.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 211 |
+
"model.layers.19.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 212 |
+
"model.layers.19.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 213 |
+
"model.layers.19.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 214 |
+
"model.layers.19.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 215 |
+
"model.layers.19.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 216 |
+
"model.layers.19.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 217 |
+
"model.layers.19.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 218 |
+
"model.layers.19.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 219 |
+
"model.layers.19.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 220 |
+
"model.layers.19.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 221 |
+
"model.layers.19.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 222 |
+
"model.layers.19.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 223 |
+
"model.layers.19.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 224 |
+
"model.layers.19.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 225 |
+
"model.layers.2.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 226 |
+
"model.layers.2.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 227 |
+
"model.layers.2.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 228 |
+
"model.layers.2.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 229 |
+
"model.layers.2.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 230 |
+
"model.layers.2.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 231 |
+
"model.layers.2.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 232 |
+
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 233 |
+
"model.layers.2.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 234 |
+
"model.layers.2.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 235 |
+
"model.layers.2.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 236 |
+
"model.layers.2.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 237 |
+
"model.layers.2.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 238 |
+
"model.layers.2.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 239 |
+
"model.layers.2.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 240 |
+
"model.layers.2.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 241 |
+
"model.layers.2.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 242 |
+
"model.layers.2.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 243 |
+
"model.layers.20.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 244 |
+
"model.layers.20.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 245 |
+
"model.layers.20.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 246 |
+
"model.layers.20.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 247 |
+
"model.layers.20.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 248 |
+
"model.layers.20.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 249 |
+
"model.layers.20.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 250 |
+
"model.layers.20.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 251 |
+
"model.layers.20.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 252 |
+
"model.layers.20.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 253 |
+
"model.layers.20.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 254 |
+
"model.layers.20.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 255 |
+
"model.layers.20.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 256 |
+
"model.layers.20.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 257 |
+
"model.layers.20.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 258 |
+
"model.layers.20.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 259 |
+
"model.layers.20.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 260 |
+
"model.layers.20.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 261 |
+
"model.layers.21.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 262 |
+
"model.layers.21.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 263 |
+
"model.layers.21.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 264 |
+
"model.layers.21.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 265 |
+
"model.layers.21.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 266 |
+
"model.layers.21.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 267 |
+
"model.layers.21.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 268 |
+
"model.layers.21.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 269 |
+
"model.layers.21.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 270 |
+
"model.layers.21.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 271 |
+
"model.layers.21.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 272 |
+
"model.layers.21.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 273 |
+
"model.layers.21.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 274 |
+
"model.layers.21.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 275 |
+
"model.layers.21.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 276 |
+
"model.layers.21.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 277 |
+
"model.layers.21.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 278 |
+
"model.layers.21.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 279 |
+
"model.layers.22.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 280 |
+
"model.layers.22.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 281 |
+
"model.layers.22.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 282 |
+
"model.layers.22.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 283 |
+
"model.layers.22.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 284 |
+
"model.layers.22.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 285 |
+
"model.layers.22.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 286 |
+
"model.layers.22.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 287 |
+
"model.layers.22.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 288 |
+
"model.layers.22.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 289 |
+
"model.layers.22.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 290 |
+
"model.layers.22.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 291 |
+
"model.layers.22.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 292 |
+
"model.layers.22.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 293 |
+
"model.layers.22.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 294 |
+
"model.layers.22.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 295 |
+
"model.layers.22.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 296 |
+
"model.layers.22.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 297 |
+
"model.layers.23.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 298 |
+
"model.layers.23.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 299 |
+
"model.layers.23.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 300 |
+
"model.layers.23.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 301 |
+
"model.layers.23.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 302 |
+
"model.layers.23.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 303 |
+
"model.layers.23.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 304 |
+
"model.layers.23.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 305 |
+
"model.layers.23.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 306 |
+
"model.layers.23.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 307 |
+
"model.layers.23.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 308 |
+
"model.layers.23.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 309 |
+
"model.layers.23.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 310 |
+
"model.layers.23.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 311 |
+
"model.layers.23.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 312 |
+
"model.layers.23.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 313 |
+
"model.layers.23.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 314 |
+
"model.layers.23.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 315 |
+
"model.layers.24.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 316 |
+
"model.layers.24.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 317 |
+
"model.layers.24.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 318 |
+
"model.layers.24.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 319 |
+
"model.layers.24.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 320 |
+
"model.layers.24.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 321 |
+
"model.layers.24.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 322 |
+
"model.layers.24.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 323 |
+
"model.layers.24.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 324 |
+
"model.layers.24.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 325 |
+
"model.layers.24.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 326 |
+
"model.layers.24.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 327 |
+
"model.layers.24.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 328 |
+
"model.layers.24.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 329 |
+
"model.layers.24.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 330 |
+
"model.layers.24.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 331 |
+
"model.layers.24.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 332 |
+
"model.layers.24.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 333 |
+
"model.layers.25.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 334 |
+
"model.layers.25.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 335 |
+
"model.layers.25.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 336 |
+
"model.layers.25.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 337 |
+
"model.layers.25.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 338 |
+
"model.layers.25.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 339 |
+
"model.layers.25.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 340 |
+
"model.layers.25.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 341 |
+
"model.layers.25.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 342 |
+
"model.layers.25.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 343 |
+
"model.layers.25.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 344 |
+
"model.layers.25.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 345 |
+
"model.layers.25.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 346 |
+
"model.layers.25.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 347 |
+
"model.layers.25.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 348 |
+
"model.layers.25.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 349 |
+
"model.layers.25.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 350 |
+
"model.layers.25.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 351 |
+
"model.layers.26.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 352 |
+
"model.layers.26.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 353 |
+
"model.layers.26.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 354 |
+
"model.layers.26.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 355 |
+
"model.layers.26.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 356 |
+
"model.layers.26.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 357 |
+
"model.layers.26.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 358 |
+
"model.layers.26.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 359 |
+
"model.layers.26.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 360 |
+
"model.layers.26.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 361 |
+
"model.layers.26.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 362 |
+
"model.layers.26.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 363 |
+
"model.layers.26.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 364 |
+
"model.layers.26.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 365 |
+
"model.layers.26.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 366 |
+
"model.layers.26.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 367 |
+
"model.layers.26.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 368 |
+
"model.layers.26.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 369 |
+
"model.layers.27.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 370 |
+
"model.layers.27.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 371 |
+
"model.layers.27.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 372 |
+
"model.layers.27.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 373 |
+
"model.layers.27.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 374 |
+
"model.layers.27.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 375 |
+
"model.layers.27.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 376 |
+
"model.layers.27.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 377 |
+
"model.layers.27.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 378 |
+
"model.layers.27.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 379 |
+
"model.layers.27.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 380 |
+
"model.layers.27.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 381 |
+
"model.layers.27.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 382 |
+
"model.layers.27.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 383 |
+
"model.layers.27.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 384 |
+
"model.layers.27.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 385 |
+
"model.layers.27.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 386 |
+
"model.layers.27.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 387 |
+
"model.layers.28.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 388 |
+
"model.layers.28.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 389 |
+
"model.layers.28.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 390 |
+
"model.layers.28.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 391 |
+
"model.layers.28.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 392 |
+
"model.layers.28.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 393 |
+
"model.layers.28.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 394 |
+
"model.layers.28.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 395 |
+
"model.layers.28.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 396 |
+
"model.layers.28.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 397 |
+
"model.layers.28.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 398 |
+
"model.layers.28.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 399 |
+
"model.layers.28.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 400 |
+
"model.layers.28.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 401 |
+
"model.layers.28.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 402 |
+
"model.layers.28.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 403 |
+
"model.layers.28.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 404 |
+
"model.layers.28.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 405 |
+
"model.layers.29.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 406 |
+
"model.layers.29.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 407 |
+
"model.layers.29.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 408 |
+
"model.layers.29.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 409 |
+
"model.layers.29.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 410 |
+
"model.layers.29.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 411 |
+
"model.layers.29.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 412 |
+
"model.layers.29.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 413 |
+
"model.layers.29.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 414 |
+
"model.layers.29.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 415 |
+
"model.layers.29.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 416 |
+
"model.layers.29.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 417 |
+
"model.layers.29.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 418 |
+
"model.layers.29.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 419 |
+
"model.layers.29.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 420 |
+
"model.layers.29.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 421 |
+
"model.layers.29.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 422 |
+
"model.layers.29.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 423 |
+
"model.layers.3.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 424 |
+
"model.layers.3.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 425 |
+
"model.layers.3.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 426 |
+
"model.layers.3.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 427 |
+
"model.layers.3.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 428 |
+
"model.layers.3.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 429 |
+
"model.layers.3.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 430 |
+
"model.layers.3.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 431 |
+
"model.layers.3.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 432 |
+
"model.layers.3.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 433 |
+
"model.layers.3.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 434 |
+
"model.layers.3.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 435 |
+
"model.layers.3.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 436 |
+
"model.layers.3.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 437 |
+
"model.layers.3.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 438 |
+
"model.layers.3.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 439 |
+
"model.layers.3.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 440 |
+
"model.layers.3.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 441 |
+
"model.layers.30.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 442 |
+
"model.layers.30.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 443 |
+
"model.layers.30.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 444 |
+
"model.layers.30.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 445 |
+
"model.layers.30.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 446 |
+
"model.layers.30.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 447 |
+
"model.layers.30.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 448 |
+
"model.layers.30.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 449 |
+
"model.layers.30.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 450 |
+
"model.layers.30.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 451 |
+
"model.layers.30.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 452 |
+
"model.layers.30.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 453 |
+
"model.layers.30.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 454 |
+
"model.layers.30.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 455 |
+
"model.layers.30.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 456 |
+
"model.layers.30.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 457 |
+
"model.layers.30.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 458 |
+
"model.layers.30.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 459 |
+
"model.layers.31.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 460 |
+
"model.layers.31.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 461 |
+
"model.layers.31.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 462 |
+
"model.layers.31.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 463 |
+
"model.layers.31.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 464 |
+
"model.layers.31.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 465 |
+
"model.layers.31.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 466 |
+
"model.layers.31.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 467 |
+
"model.layers.31.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 468 |
+
"model.layers.31.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 469 |
+
"model.layers.31.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 470 |
+
"model.layers.31.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 471 |
+
"model.layers.31.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 472 |
+
"model.layers.31.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 473 |
+
"model.layers.31.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 474 |
+
"model.layers.31.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 475 |
+
"model.layers.31.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 476 |
+
"model.layers.31.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 477 |
+
"model.layers.32.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 478 |
+
"model.layers.32.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 479 |
+
"model.layers.32.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 480 |
+
"model.layers.32.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 481 |
+
"model.layers.32.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 482 |
+
"model.layers.32.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 483 |
+
"model.layers.32.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 484 |
+
"model.layers.32.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 485 |
+
"model.layers.32.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 486 |
+
"model.layers.32.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 487 |
+
"model.layers.32.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 488 |
+
"model.layers.32.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 489 |
+
"model.layers.32.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 490 |
+
"model.layers.32.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 491 |
+
"model.layers.32.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 492 |
+
"model.layers.32.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 493 |
+
"model.layers.32.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 494 |
+
"model.layers.32.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 495 |
+
"model.layers.33.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 496 |
+
"model.layers.33.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 497 |
+
"model.layers.33.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 498 |
+
"model.layers.33.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 499 |
+
"model.layers.33.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 500 |
+
"model.layers.33.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 501 |
+
"model.layers.33.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 502 |
+
"model.layers.33.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 503 |
+
"model.layers.33.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 504 |
+
"model.layers.33.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 505 |
+
"model.layers.33.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 506 |
+
"model.layers.33.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 507 |
+
"model.layers.33.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 508 |
+
"model.layers.33.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 509 |
+
"model.layers.33.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 510 |
+
"model.layers.33.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 511 |
+
"model.layers.33.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 512 |
+
"model.layers.33.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 513 |
+
"model.layers.34.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 514 |
+
"model.layers.34.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 515 |
+
"model.layers.34.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 516 |
+
"model.layers.34.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 517 |
+
"model.layers.34.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 518 |
+
"model.layers.34.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 519 |
+
"model.layers.34.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 520 |
+
"model.layers.34.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 521 |
+
"model.layers.34.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 522 |
+
"model.layers.34.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 523 |
+
"model.layers.34.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 524 |
+
"model.layers.34.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 525 |
+
"model.layers.34.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 526 |
+
"model.layers.34.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 527 |
+
"model.layers.34.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 528 |
+
"model.layers.34.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 529 |
+
"model.layers.34.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 530 |
+
"model.layers.34.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 531 |
+
"model.layers.35.input_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 532 |
+
"model.layers.35.mlp.down_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 533 |
+
"model.layers.35.mlp.down_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 534 |
+
"model.layers.35.mlp.gate_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 535 |
+
"model.layers.35.mlp.gate_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 536 |
+
"model.layers.35.mlp.up_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 537 |
+
"model.layers.35.mlp.up_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 538 |
+
"model.layers.35.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
|
| 539 |
+
"model.layers.35.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
|
| 540 |
+
"model.layers.35.self_attn.k_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 541 |
+
"model.layers.35.self_attn.k_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 542 |
+
"model.layers.35.self_attn.o_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 543 |
+
"model.layers.35.self_attn.o_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 544 |
+
"model.layers.35.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
|
| 545 |
+
"model.layers.35.self_attn.q_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 546 |
+
"model.layers.35.self_attn.q_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 547 |
+
"model.layers.35.self_attn.v_proj.packed_weight_indices": "model-00002-of-00002.safetensors",
|
| 548 |
+
"model.layers.35.self_attn.v_proj.row_norms": "model-00002-of-00002.safetensors",
|
| 549 |
+
"model.layers.4.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 550 |
+
"model.layers.4.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 551 |
+
"model.layers.4.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 552 |
+
"model.layers.4.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 553 |
+
"model.layers.4.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 554 |
+
"model.layers.4.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 555 |
+
"model.layers.4.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 556 |
+
"model.layers.4.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 557 |
+
"model.layers.4.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 558 |
+
"model.layers.4.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 559 |
+
"model.layers.4.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 560 |
+
"model.layers.4.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 561 |
+
"model.layers.4.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 562 |
+
"model.layers.4.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 563 |
+
"model.layers.4.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 564 |
+
"model.layers.4.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 565 |
+
"model.layers.4.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 566 |
+
"model.layers.4.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 567 |
+
"model.layers.5.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 568 |
+
"model.layers.5.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 569 |
+
"model.layers.5.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 570 |
+
"model.layers.5.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 571 |
+
"model.layers.5.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 572 |
+
"model.layers.5.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 573 |
+
"model.layers.5.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 574 |
+
"model.layers.5.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 575 |
+
"model.layers.5.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 576 |
+
"model.layers.5.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 577 |
+
"model.layers.5.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 578 |
+
"model.layers.5.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 579 |
+
"model.layers.5.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 580 |
+
"model.layers.5.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 581 |
+
"model.layers.5.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 582 |
+
"model.layers.5.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 583 |
+
"model.layers.5.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 584 |
+
"model.layers.5.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 585 |
+
"model.layers.6.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 586 |
+
"model.layers.6.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 587 |
+
"model.layers.6.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 588 |
+
"model.layers.6.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 589 |
+
"model.layers.6.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 590 |
+
"model.layers.6.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 591 |
+
"model.layers.6.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 592 |
+
"model.layers.6.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 593 |
+
"model.layers.6.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 594 |
+
"model.layers.6.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 595 |
+
"model.layers.6.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 596 |
+
"model.layers.6.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 597 |
+
"model.layers.6.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 598 |
+
"model.layers.6.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 599 |
+
"model.layers.6.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 600 |
+
"model.layers.6.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 601 |
+
"model.layers.6.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 602 |
+
"model.layers.6.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 603 |
+
"model.layers.7.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 604 |
+
"model.layers.7.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 605 |
+
"model.layers.7.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 606 |
+
"model.layers.7.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 607 |
+
"model.layers.7.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 608 |
+
"model.layers.7.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 609 |
+
"model.layers.7.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 610 |
+
"model.layers.7.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 611 |
+
"model.layers.7.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 612 |
+
"model.layers.7.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 613 |
+
"model.layers.7.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 614 |
+
"model.layers.7.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 615 |
+
"model.layers.7.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 616 |
+
"model.layers.7.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 617 |
+
"model.layers.7.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 618 |
+
"model.layers.7.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 619 |
+
"model.layers.7.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 620 |
+
"model.layers.7.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 621 |
+
"model.layers.8.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 622 |
+
"model.layers.8.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 623 |
+
"model.layers.8.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 624 |
+
"model.layers.8.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 625 |
+
"model.layers.8.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 626 |
+
"model.layers.8.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 627 |
+
"model.layers.8.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 628 |
+
"model.layers.8.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 629 |
+
"model.layers.8.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 630 |
+
"model.layers.8.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 631 |
+
"model.layers.8.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 632 |
+
"model.layers.8.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 633 |
+
"model.layers.8.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 634 |
+
"model.layers.8.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 635 |
+
"model.layers.8.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 636 |
+
"model.layers.8.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 637 |
+
"model.layers.8.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 638 |
+
"model.layers.8.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 639 |
+
"model.layers.9.input_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 640 |
+
"model.layers.9.mlp.down_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 641 |
+
"model.layers.9.mlp.down_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 642 |
+
"model.layers.9.mlp.gate_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 643 |
+
"model.layers.9.mlp.gate_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 644 |
+
"model.layers.9.mlp.up_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 645 |
+
"model.layers.9.mlp.up_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 646 |
+
"model.layers.9.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
|
| 647 |
+
"model.layers.9.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
|
| 648 |
+
"model.layers.9.self_attn.k_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 649 |
+
"model.layers.9.self_attn.k_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 650 |
+
"model.layers.9.self_attn.o_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 651 |
+
"model.layers.9.self_attn.o_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 652 |
+
"model.layers.9.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
|
| 653 |
+
"model.layers.9.self_attn.q_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 654 |
+
"model.layers.9.self_attn.q_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 655 |
+
"model.layers.9.self_attn.v_proj.packed_weight_indices": "model-00001-of-00002.safetensors",
|
| 656 |
+
"model.layers.9.self_attn.v_proj.row_norms": "model-00001-of-00002.safetensors",
|
| 657 |
+
"model.norm.weight": "model-00002-of-00002.safetensors"
|
| 658 |
+
}
|
| 659 |
+
}
|
tokenizer/chat_template.jinja
ADDED
|
@@ -0,0 +1,89 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if message.content is string %}
|
| 27 |
+
{%- set content = message.content %}
|
| 28 |
+
{%- else %}
|
| 29 |
+
{%- set content = '' %}
|
| 30 |
+
{%- endif %}
|
| 31 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 32 |
+
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
| 33 |
+
{%- elif message.role == "assistant" %}
|
| 34 |
+
{%- set reasoning_content = '' %}
|
| 35 |
+
{%- if message.reasoning_content is string %}
|
| 36 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 37 |
+
{%- else %}
|
| 38 |
+
{%- if '</think>' in content %}
|
| 39 |
+
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 40 |
+
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
| 41 |
+
{%- endif %}
|
| 42 |
+
{%- endif %}
|
| 43 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 44 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 45 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 46 |
+
{%- else %}
|
| 47 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 48 |
+
{%- endif %}
|
| 49 |
+
{%- else %}
|
| 50 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 51 |
+
{%- endif %}
|
| 52 |
+
{%- if message.tool_calls %}
|
| 53 |
+
{%- for tool_call in message.tool_calls %}
|
| 54 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 55 |
+
{{- '\n' }}
|
| 56 |
+
{%- endif %}
|
| 57 |
+
{%- if tool_call.function %}
|
| 58 |
+
{%- set tool_call = tool_call.function %}
|
| 59 |
+
{%- endif %}
|
| 60 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 61 |
+
{{- tool_call.name }}
|
| 62 |
+
{{- '", "arguments": ' }}
|
| 63 |
+
{%- if tool_call.arguments is string %}
|
| 64 |
+
{{- tool_call.arguments }}
|
| 65 |
+
{%- else %}
|
| 66 |
+
{{- tool_call.arguments | tojson }}
|
| 67 |
+
{%- endif %}
|
| 68 |
+
{{- '}\n</tool_call>' }}
|
| 69 |
+
{%- endfor %}
|
| 70 |
+
{%- endif %}
|
| 71 |
+
{{- '<|im_end|>\n' }}
|
| 72 |
+
{%- elif message.role == "tool" %}
|
| 73 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 74 |
+
{{- '<|im_start|>user' }}
|
| 75 |
+
{%- endif %}
|
| 76 |
+
{{- '\n<tool_response>\n' }}
|
| 77 |
+
{{- content }}
|
| 78 |
+
{{- '\n</tool_response>' }}
|
| 79 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 80 |
+
{{- '<|im_end|>\n' }}
|
| 81 |
+
{%- endif %}
|
| 82 |
+
{%- endif %}
|
| 83 |
+
{%- endfor %}
|
| 84 |
+
{%- if add_generation_prompt %}
|
| 85 |
+
{{- '<|im_start|>assistant\n' }}
|
| 86 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 87 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 88 |
+
{%- endif %}
|
| 89 |
+
{%- endif %}
|
tokenizer/tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506
|
| 3 |
+
size 11422650
|
tokenizer/tokenizer_config.json
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|im_end|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"is_local": true,
|
| 9 |
+
"local_files_only": false,
|
| 10 |
+
"model_max_length": 131072,
|
| 11 |
+
"pad_token": "<|endoftext|>",
|
| 12 |
+
"split_special_tokens": false,
|
| 13 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 14 |
+
"unk_token": null
|
| 15 |
+
}
|
transformer/config.json
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_class_name": "Flux2Transformer2DModel",
|
| 3 |
+
"_diffusers_version": "0.39.0",
|
| 4 |
+
"_name_or_path": "/opt/oq9b/hf/hub/models--black-forest-labs--FLUX.2-klein-9B/snapshots/92196c8e11f7b6cf2b7493e037d8c5345c559216/transformer",
|
| 5 |
+
"attention_head_dim": 128,
|
| 6 |
+
"axes_dims_rope": [
|
| 7 |
+
32,
|
| 8 |
+
32,
|
| 9 |
+
32,
|
| 10 |
+
32
|
| 11 |
+
],
|
| 12 |
+
"eps": 1e-06,
|
| 13 |
+
"guidance_embeds": false,
|
| 14 |
+
"in_channels": 128,
|
| 15 |
+
"joint_attention_dim": 12288,
|
| 16 |
+
"mlp_ratio": 3.0,
|
| 17 |
+
"num_attention_heads": 32,
|
| 18 |
+
"num_layers": 8,
|
| 19 |
+
"num_single_layers": 24,
|
| 20 |
+
"out_channels": null,
|
| 21 |
+
"patch_size": 1,
|
| 22 |
+
"quantization_config": {
|
| 23 |
+
"activation_bits": 4,
|
| 24 |
+
"activation_eps": 1e-10,
|
| 25 |
+
"activation_kernel_backend": "triton_cuda",
|
| 26 |
+
"activation_norm_dtype": "float32",
|
| 27 |
+
"adaln_group_size": 64,
|
| 28 |
+
"adaln_policy": "int4_rtn",
|
| 29 |
+
"artifact_format_version": 1,
|
| 30 |
+
"block_size": "paper",
|
| 31 |
+
"codebook": "lloyd_max",
|
| 32 |
+
"codebook_dtype": "float32",
|
| 33 |
+
"codebook_version": 2,
|
| 34 |
+
"modules_dtype_dict": {},
|
| 35 |
+
"modules_to_convert": [],
|
| 36 |
+
"modules_to_not_convert": [],
|
| 37 |
+
"modules_to_use_adaln": [],
|
| 38 |
+
"packed_matmul_block_k": 128,
|
| 39 |
+
"packed_matmul_block_m": 64,
|
| 40 |
+
"packed_matmul_block_n": 64,
|
| 41 |
+
"packed_matmul_num_warps": 4,
|
| 42 |
+
"quant_method": "orbitquant",
|
| 43 |
+
"rotation": "rpbh",
|
| 44 |
+
"rotation_seed": 0,
|
| 45 |
+
"row_norm_dtype": "bfloat16",
|
| 46 |
+
"runtime_mode": "auto_fused",
|
| 47 |
+
"target_policy": "auto",
|
| 48 |
+
"weight_bits": 4,
|
| 49 |
+
"weight_pack_dtype": "uint8"
|
| 50 |
+
},
|
| 51 |
+
"rope_theta": 2000,
|
| 52 |
+
"timestep_guidance_channels": 256
|
| 53 |
+
}
|
transformer/diffusion_pytorch_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c5eaeecc5c474756a28de32d67edd8320ac44fcdb9fc5f71d83ab4cdeccac84e
|
| 3 |
+
size 4704766296
|
vae/config.json
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_class_name": "AutoencoderKLFlux2",
|
| 3 |
+
"_diffusers_version": "0.39.0",
|
| 4 |
+
"_name_or_path": "/opt/oq9b/hf/hub/models--black-forest-labs--FLUX.2-klein-9B/snapshots/92196c8e11f7b6cf2b7493e037d8c5345c559216/vae",
|
| 5 |
+
"act_fn": "silu",
|
| 6 |
+
"batch_norm_eps": 0.0001,
|
| 7 |
+
"batch_norm_momentum": 0.1,
|
| 8 |
+
"block_out_channels": [
|
| 9 |
+
128,
|
| 10 |
+
256,
|
| 11 |
+
512,
|
| 12 |
+
512
|
| 13 |
+
],
|
| 14 |
+
"decoder_block_out_channels": null,
|
| 15 |
+
"down_block_types": [
|
| 16 |
+
"DownEncoderBlock2D",
|
| 17 |
+
"DownEncoderBlock2D",
|
| 18 |
+
"DownEncoderBlock2D",
|
| 19 |
+
"DownEncoderBlock2D"
|
| 20 |
+
],
|
| 21 |
+
"force_upcast": true,
|
| 22 |
+
"in_channels": 3,
|
| 23 |
+
"latent_channels": 32,
|
| 24 |
+
"layers_per_block": 2,
|
| 25 |
+
"mid_block_add_attention": true,
|
| 26 |
+
"norm_num_groups": 32,
|
| 27 |
+
"out_channels": 3,
|
| 28 |
+
"patch_size": [
|
| 29 |
+
2,
|
| 30 |
+
2
|
| 31 |
+
],
|
| 32 |
+
"sample_size": 1024,
|
| 33 |
+
"up_block_types": [
|
| 34 |
+
"UpDecoderBlock2D",
|
| 35 |
+
"UpDecoderBlock2D",
|
| 36 |
+
"UpDecoderBlock2D",
|
| 37 |
+
"UpDecoderBlock2D"
|
| 38 |
+
],
|
| 39 |
+
"use_post_quant_conv": true,
|
| 40 |
+
"use_quant_conv": true
|
| 41 |
+
}
|
vae/diffusion_pytorch_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ca70d2202afe6415bdbcb8793ba8cd99fd159cfe6192381504d6c4d3036e0f04
|
| 3 |
+
size 168120878
|