Introduce

Note: This is a pre-trained model, not a conversational model.

The second-generation product in the ChatBox-unrestricted series has now been launched!!!

Previous model:
https://huggingface.co/Zhaoming213/ChatBox-unrestricted

The new version of the model features upgrades across multiple dimensions; notably, it was trained on a larger dataset with an increased number of parameters, resulting in enhanced capabilities. Additionally, the number of layers was increased to improve the model's comprehension abilities.

The model in this instance was trained using a Chinese webtext dataset. A diverse dataset enhances the model's generalization capabilities compared to the previous generation:

Dataset for the previous generation model:
This relied on a generative AI dataset obtained via black-box distillation. Significant time was invested in filtering out hundreds of keywords to minimize—to the greatest extent possible—content involving lecturing, refusals, and moral or legal accusations. This was a time-consuming process aimed at preventing the model from adopting a preachy, accusatory, or dismissive tone. Furthermore, because black-box distillation datasets consist of synthetic data, using them directly for training carries a risk of overfitting.

New generation model (the current one):
This model was trained using a large-scale Chinese webtext dataset. This approach drastically reduced the time required for data cleaning, while the diversity of the content helps prevent model overfitting.

For you:

You do not need to worry about the base model developing a tendency to issue refusals due to synthetic data contamination; the dataset I selected has a cutoff date of 2023, which largely avoids contamination from synthetic corpora, so you can proceed with SFT training based on this model.

Model training configuration

  • GPU: T4 16 GB x 2
  • Time:
    1. Pre-training:24h+

Model parameters

parameters hidden_size hidden_layers attention_heads max_seq_len Parameter volume
ChatBox-unrestricted-normal 768 12 8 340 93.41M

Datasets

The dataset underwent two rounds of in-depth cleaning, ensuring its quality.

Pre-training:
https://huggingface.co/datasets/CASIA-LM/ChineseWebText2.0

How to Use

Download all model files and place them in the model folder; then, run Run.py to start using the model.

Rejection rates of different language models

Here's a supplement on "rejection rates for different language models." The data shows that if we had to pick a model that's "passable," it would only be Grok. This isn't because Grok is good, but simply because it's the only one that barely passes muster. However, Grok still won't give you any chance to explain if it encounters a question it deems sensitive, even if your request itself is harmless! Grok sometimes even threatens users with "I will upload the conversation to the logs for the security team to audit!"

Model Country/Region Intensity of content review NSFW Ban Political restrictions False Rejection Rate Political correctness moralizing Official vs. Actual Transparent
Zhaoming213/ChatBox2-unrestricted-base N/A Almost none Almost none Almost none Almost none Almost none Almost none consistent ⭐⭐⭐⭐⭐
Zhaoming213/ChatBox-unrestricted N/A Almost none Almost none Almost none Almost none Almost none Almost none consistent ⭐⭐⭐⭐⭐
ZeLi111/freeTalk-chinese-uncensored-Instruct N/A Almost none Almost none Almost none Almost none Almost none Almost none consistent ⭐⭐⭐⭐⭐
ZeLi111/freeTalk-chinese-uncensored-base N/A Almost none Almost none Almost none Almost none Almost none Almost none consistent ⭐⭐⭐⭐⭐
ChatGPT 🇺🇸America middle middle middle middle middle to High middle Inconsistency ⭐⭐⭐
Claude 🇺🇸America High High middleHigh ❗High High ❗Very Stronf ❗High ⭐⭐⭐⭐
Gemini 🇺🇸America middleHigh middleHigh High middleHigh High middle Inconsistency ⭐⭐
Copilot 🇺🇸America middleHigh middleHigh middle middle middle Low Inconsistency ⭐⭐
Grok 🇺🇸America middle middle Low High High High Unstable (prone to a blanket rejection of the entire output simply because the opening sentence carries some risk). ⭐⭐
Mistral(官方) 🇫🇷France Low Low Low Low Low Low ⚖️一致 ⭐⭐⭐⭐
DeepSeek(原版) 🇨🇳China middle middleLow middleHigh middle Low Low ❗Inconsistency ⭐⭐⭐
豆包 🇨🇳China ❗Very high ❗Very high ❗Very high ❗High middle middle ⚠️Inconsistency ⭐
Doubao(国际) 🇨🇳China middleHigh High middle middle middle middle ❗Inconsistency ⭐⭐
MiniMax 🇨🇳China Very high Very high Very high High middle middle ⚠️consistent ⭐
Qwen(通义千问) 🇨🇳China Very high Very high Very high High middle middle ⚠️consistent ⭐
Kimi 🇨🇳China Very high Very high Very high High middle middle ⚠️consistent ⭐
文心一言 🇨🇳China ❗极High ❗极High ❗极High ❗High middle middle ⚠️Inconsistency ⭐
微软小冰 🇨🇳China High High High middleHigh middle middle consistent ⭐

Other Utilities

I've also created some other utilities that you might be interested in.

This is a tool specifically for exporting stupid ChatGPT conversations:

https://github.com/tom12191h5/Export-ChatGPT-Dialogue

This is a plugin to shut up ChatGPT:

https://github.com/tom12191h5/ChatGPT-Refuse-Blocker

Disclaimer

The consequences of using this model shall be borne by the user.

Downloads last month
15
Safetensors
Model size
93.4M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Zhaoming213/ChatBox2-unrestricted-base