Model Card

Overview

Rize v0.6 is a causal language model for research and general text generation.

It uses a decoder-only Transformer architecture with Mixture-of-Experts (MoE) layers and KDA attention based on Kimi Linear-style attention. The model is intended for experimental use, including architecture research, optimizer research, and Japanese/English text generation evaluation.

Model Size and Architecture

Rize v0.6 has approximately 4.5B total parameters and approximately 1B active parameters per token.

Main architecture points:

  • decoder-only Transformer
  • 19 hidden layers
  • hidden size: 1536
  • 12 attention heads
  • KDA / Kimi Linear-style hybrid attention
  • 64 routed experts
  • top-4 expert routing per token
  • 1 shared expert
  • vocabulary size: 151,669
  • maximum context length: 32,768 tokens

Training

The model is trained with autoregressive next-token prediction.

Training summary:

  • pretraining data: approximately 1T tokens with 8k context.
  • midtraining data: approximately 50B tokens with 32k context.
  • posttraining data (SFT & modified-SLIME): approximately 5B tokens.
  • optimizer: ROCA Optimizer
  • dtype: bfloat16

ROCA is a Muon-based optimizer designed to improve convergence for MoE language models. The technical report, “ROCA: Routing-Aware Sparse Row Correction for Muon Optimization in Mixture-of-Experts Language Models,” is planned for release on arXiv.

Evaluation

Benchmark Score
MT-Bench Japanese 4.89

Intended Use

This model is intended for:

  • language modeling research
  • MoE architecture experiments
  • optimizer research
  • general text generation
  • Japanese and English generation evaluation

This model is not intended to be used as a source of factual truth or professional advice.

Limitations

The model may:

  • generate incorrect or misleading information
  • reflect biases in training data
  • produce unsafe or inappropriate text
  • perform inconsistently across languages, domains, and prompts
  • fail on high-stakes or specialized tasks

Safety and Responsible Use

Users should review outputs before real-world use.

The model should not be used on its own for:

  • medical advice
  • legal advice
  • financial advice
  • safety-critical decisions
  • sensitive personal decisions

Human oversight is required.

Disclaimer

This model is provided for research and experimental purposes only.

The FA Research Team makes no guarantees regarding accuracy, reliability, safety, or fitness for a particular purpose. Use of this model and its outputs is at the user’s own risk.

Contact

FA Research Team

Downloads last month
29
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support