BananaMind-2-Pro-Thinking
BananaMind-2-Pro-Thinking is a fine-tuned version of the BananaMind-2-Pro Base model using our proprietary datasets and training pipeline.
The original BananaMind-2-Pro model was not designed as a thinking model. Through our continued pretraining (CPT) and supervised fine-tuning (SFT) pipeline, we were able to induce substantially stronger reasoning behavior while retaining the underlying architecture.
Credit for the original model and architecture belongs to BananaMind. This repository contains our post-trained checkpoint and does not claim authorship of the original model.
Model Details
| Property | Value |
|---|---|
| Architecture | BananaMind2Pro decoder-only Transformer |
| Parameters | ~138M |
| Context Length | 3,072 tokens |
| Base Model | BananaMind-2-Pro |
| Training Type | CPT + SFT |
| Model Type | Autoregressive causal language model |
| Release Type | Research checkpoint |
Training
Training was performed on a single NVIDIA RTX 3090.
| Training Parameter | Value |
|---|---|
| Sequence Length | 3,072 |
| Batch Size | 4 |
| Gradient Accumulation | 4 |
| Effective Batch Size | 16 sequences |
| Peak Learning Rate | 1e-4 |
Training data consisted of proprietary datasets used for both continued pretraining and supervised fine-tuning.
Research Motivation
This checkpoint was created primarily as a research experiment to evaluate our post-training pipeline on an existing model in a similar parameter class.
Our previous experiments using the same general approach on our custom Sabaki-Preview model did not produce the results we expected. BananaMind-2-Pro served as a control for determining whether the limitation was primarily caused by our post-training methodology or by characteristics of the underlying Sabaki model.
The results suggest that the CPT + SFT pipeline is capable of inducing meaningful reasoning improvements in a model of this size.
This model should therefore be considered a research checkpoint rather than a production model.
Performance
We evaluated both the original BananaMind-2-Pro Base checkpoint and our BananaMind-2-Pro-Thinking checkpoint using the same evaluation setup.
| Benchmark | Base | Thinking | Change |
|---|---|---|---|
| ARC Easy | 55.26% | 58.80% | +3.54 |
| ARC Challenge | 27.82% | 29.61% | +1.79 |
| PIQA | 67.52% | 68.28% | +0.76 |
| HellaSwag* | 42.78% | 44.70% | +1.92 |
| ArithMark 3 | 38.20% | 38.70% | +0.50 |
* HellaSwag was evaluated on a 1,000-question subset rather than the complete benchmark.
The Thinking checkpoint improved over the Base model on all five reported evaluations, with the largest absolute gain appearing on ARC Easy.
Open SLM Leaderboard
Based on these benchmark results and the scoring weighting used by the AxiomicLabs Open SLM Leaderboard, our scores would place BananaMind-2-Pro-Thinking at #1 on the leaderboard as of September 27, 2026.
However, we do not consider this an official leaderboard result.
We intentionally do not claim an official ranking because:
- HellaSwag was evaluated using only 1,000 questions, rather than the complete benchmark.
- The reported evaluations have not yet received independent third-party verification.
- Evaluation implementation differences can affect results even when benchmark names are identical.
For those reasons, we believe submitting these numbers as directly equivalent to independently evaluated full-benchmark results would be unfair.
We encourage independent reproduction and verification of these results.
Thinking Behavior
The original BananaMind-2-Pro Base model was not explicitly trained as a reasoning or "thinking" model.
Our CPT + SFT pipeline changes the model's behavior so that it can produce more structured intermediate reasoning before arriving at a final response.
The model remains small at approximately 138M parameters, so users should not expect reasoning performance comparable to substantially larger language models. This release is primarily intended to study how much reasoning behavior can be induced through post-training at this scale.
Intended Use
BananaMind-2-Pro-Thinking is intended for:
- Research into reasoning behavior in small language models
- Evaluation of continued-pretraining and SFT techniques
- Comparison against the original BananaMind-2-Pro checkpoint
- Small-model reasoning experiments
- Research into compute-efficient post-training
- Reproduction and independent benchmarking
It is not intended for high-stakes or production-critical applications without additional evaluation.
Limitations
This model inherits limitations from the original BananaMind-2-Pro model and may introduce additional behaviors as a consequence of our post-training data.
Known limitations include:
- The model is only approximately 138M parameters and has limited world knowledge and capacity.
- Reasoning traces may contain incorrect intermediate steps even when the final answer appears plausible.
- The model may hallucinate facts or produce confidently incorrect responses.
- Benchmark improvements do not necessarily translate directly to all real-world tasks.
- HellaSwag results currently use a 1,000-question subset and should not be interpreted as equivalent to a full-benchmark evaluation.
- Results have not yet been independently verified by a third party.
- The proprietary training dataset is not included with this release.
Users should independently evaluate the model for their intended application.
License
This model is a derivative of BananaMind/BananaMind-2-Pro and is distributed under the BananaMind Community License 1.0.
Use, modification, redistribution, and commercial deployment of this model are subject to the terms of that license.
The original BananaMind-2-Pro repository notes that commercial products or services exceeding the thresholds defined in Section 1 of the BananaMind Community License 1.0 require a separate commercial license from Banaxi-Tech.
Users of this derivative model are responsible for reviewing and complying with the full license terms provided with the original BananaMind-2-Pro release.
Attribution
The original BananaMind-2-Pro architecture and model were created by BananaMind.
This release applies our proprietary continued-pretraining and supervised fine-tuning datasets and pipeline to that model.
The original creator's copyright and license notices are retained in this release. When redistributing this model or derivatives, please preserve all applicable notices and provide appropriate credit to BananaMind.
Our modifications and training should not be interpreted as ownership of the original BananaMind model or architecture.
License bananamind-community-license-1.0
Acknowledgements
Thanks to BananaMind for creating and releasing the original BananaMind-2-Pro model.
- Downloads last month
- 372
Model tree for Local-Axiom-AI/BananaMind-2-Pro-Thinking
Base model
BananaMind/BananaMind-2-Pro