Léon Simmons
Avicennasis
AI & ML interests
None yet
Recent Activity
new activity about 6 hours ago
autotrust/GEV-26B-Decide:MLX port for Apple silicon (System 1): Avicennasis/GEV-26B-Decide-mlx-8bit / -4bit updated a model 1 day ago
Avicennasis/GEV-26B-Decide-mlx-8bit published a model 1 day ago
Avicennasis/GEV-26B-Decide-mlx-8bitOrganizations
Fix: add the missing @capture_outputs decorator (output_hidden_states is None) Fixes the issue reported in discussion #6. `K2HorizonModel.forward` carries `@auto_docstring` but not `@capture_outputs`, so `output_hidden_states=True` returns `None` (any tool that reads the residual stream per layer fails). Adding the decorator plus the import makes `outputs.hidden_states` a tuple of len `num_hidden_layers + 1`; greedy decoding is unchanged. Verified on 0.9B: 29 states (28 layers + 1), `[0]` == embeddings and `[-1]` == `last_hidden_state`; generation byte-identical with and without the decorator. The change is the same for every full-model repo in the family (the file is byte-identical across sizes).
#5 opened 4 days ago
by
Avicennasis
Fix: add the missing @capture_outputs decorator (output_hidden_states is None) Fixes the issue reported in discussion #6. `K2HorizonModel.forward` carries `@auto_docstring` but not `@capture_outputs`, so `output_hidden_states=True` returns `None` (any tool that reads the residual stream per layer fails). Adding the decorator plus the import makes `outputs.hidden_states` a tuple of len `num_hidden_layers + 1`; greedy decoding is unchanged. Verified on 0.9B: 29 states (28 layers + 1), `[0]` == embeddings and `[-1]` == `last_hidden_state`; generation byte-identical with and without the decorator. The change is the same for every full-model repo in the family (the file is byte-identical across sizes).
#6 opened 4 days ago
by
Avicennasis
Fix: add the missing @capture_outputs decorator (output_hidden_states is None) Fixes the issue reported in discussion #6. `K2HorizonModel.forward` carries `@auto_docstring` but not `@capture_outputs`, so `output_hidden_states=True` returns `None` (any tool that reads the residual stream per layer fails). Adding the decorator plus the import makes `outputs.hidden_states` a tuple of len `num_hidden_layers + 1`; greedy decoding is unchanged. Verified on 0.9B: 29 states (28 layers + 1), `[0]` == embeddings and `[-1]` == `last_hidden_state`; generation byte-identical with and without the decorator. The change is the same for every full-model repo in the family (the file is byte-identical across sizes).
#9 opened 4 days ago
by
Avicennasis