AI & ML interests

Computer Vision Technology and Data Collection for Anime Waifu

Recent Activity

AbstractPhil 
posted an update 3 days ago
view post
Post
89
GPT, Gemini, Claude, and I have identified a multitude of direct utilities for Beatrix useful for diffusion conditioning in very powerful geometric formats. We have also identified multiple weaknesses to compensate for, multiple strengths to augment, the cause of the final layer's weak erank output state, and an emergent mathematical property of calculation in this format. The final stage directional magnitude is overwhelming and becoming amplitude.

There is a full article brewing for this information, including a massive set of information already learned from Beatrix V3 that could not be extracted from the 2s variant.

As the model trains, the amplitude begins to strengthen over and over. The weak tokenization processing from splat attention, forms the internal state of the model towards a bloating fashion. This is due to the articulation applied by the structure of the aleph addressing.

This creates massive erank geometry naturally, exhausting the space, producing comprehensively complex geometric structures. This internal structure here is weakly bound to the internal bytes, causing recon to weaken over time >2048, producing the output tokenization to be weaker at higher token lengths. Training improves this but is not known to solve it.

Along this chain the final layer has formed a sort of unexpected behavior, an amplitude behavior. I've met amplitude responses before in multiple models, and even attempted to curated magnitude through flow matching to some success, however amplitude in that nature is costly and adds additional overhead to the train so I'll need to come up with something more careful, and potentially something more clever than just attaching a composite or an energy dampener.

Attention will be solved by introducing various MHA layers throughout, ensuring the recon through the depth of the model survives. With that we'll want to ensure large erank composites form as well, allowing those humongous geometric structures to form and contribute.
  • 3 replies
·
AbstractPhil 
posted an update 7 days ago
view post
Post
64
My apologies for the incorrect format for the AMOE arms from the experimental branch. They have been saving as torch objects. They are now correctly saving as safetensors format. My apologies for the inconvenience this may cause for you use. I will be modifying the codespaces to use the correct safetensors formats.

After the first 20.9b tokens trained, the real experiments begins. Beatrix V3's first prepped-state modular command structure has been attached for dynamic training. These arms will exist as appendages for Beatrix - trained alongside with the trunk until the end of the run.

These exist for experimental extraction, analysis, distillation experiments, memory experiments, mathematics experiments, and more. Each arm will be built along the chain for specific test cases. Expectation for each is already lined up and the outcomes are tested for, but the model still may face instability and must be monitored.

As her first arm learns tinystories, she builds direct composite semantic structure throughout this system. Think of it like, the first higher-functioning cognition attachment.

She's still very naïve and structurally unaware, so attaching new limbs is essentially extending a structure that is not yet finished forming. Nothing but fragments of issued information from an unknown source.

In this case, this structure has been tested hundreds of times to ensure she will not simply collapse during training by having this attached.
  • 4 replies
·
prithivMLmods 
posted an update 8 days ago
view post
Post
3349
Qwen-Image-2.1 Plug and Play LoRA App is now live on Hugging Face Spaces.

🔗 Space: prithivMLmods/Qwen-Image-2.1-LoRAs-PnP

It supports standard inference, 4-step Turbo inference, custom LoRA lazy repacks, and LoRA Plug and Play (PnP), all in one setting!

🔗 Qwen-Image-2.1 Image-to-Image LoRAs: https://huggingface.co/collections/prithivMLmods/qwen-image-21-image-to-image-loras

🔗 GitHub: https://github.com/PRITHIVSAKTHIUR/Qwen-Image-2.1-LoRAs-PnP

To learn more, visit the app page or the respective model pages.
AbstractPhil 
posted an update 11 days ago
view post
Post
101
12 day cook for mini beatrix v3 begins. This model's byte input is formatted using a method dubbed atlas input.

ETA OCTOBER 2 2026

AbstractPhil/alephllm-mini-beatrix-training

https://github.com/AbstractEyes/geolip-bytelex
https://github.com/AbstractEyes/alephllm

Upgrades:
* 64 billion byte training pipeline up from 16 billion
* 32 block depth 376.0M in v3 up from 20 block 237.1M in 2s.
* Active aleph head, repaired via the 2s faults and a large series of tests.
* Byte atlas gateway router, explained below.
* Guaranteed convergence follow-up AMOE arms on pretrain, fused into the final form, trained together over time to increase the collective capacity.
* Multi-tokenizer oriented post-training arms distilled from multiple experts; E.G. Qwen 3.8 27b multi-layer teacher/student arms, CLIP big_g, Bert Code, and more.
* Special token word implementation via AMOE arms is now tested up to 240 special tokens for routing. Theoretically each can implement it's own sub-arm aka nested commands. E.G; <think><think_symbolic> ... </think_symbolic></think>
* Fused words post-training for faster inference.

# The Atlas
This atlas structure contains the conjoined shape of 12 tokenizers represented in the trigram format. This is used to predict difficulty in the overlaps, as per determined by the average byte overlap measured via corpus text and the compared overlap. This accuracy is only related to difficulty but it provides pre-training difficulty assessment that we will use to test post-training accuracy with it. This will determine if we can precalculate the likelihood of byte difficulty via tokenizer shape in byte form, for the multibyte fusion upcoming arm experiments for v3.

The reason for this, is distillation. We need to train Beatrix to behave with multiple tokenizers, and this theory is showing accuracy with v1 and v2, but the 32 block depth of v3 will answer many questions alongside of the structure.
  • 4 replies
·
prithivMLmods 
posted an update 16 days ago
view post
Post
818
VisionGuardrail EVO-2, a multimodal image-classification content-safety model based on Qwen/Qwen3.8-27B, is now available on the Hub!

Stricter image classification than before, with a dense 27-billion-parameter multimodal model, more precise reasoning, and improved captions for classifying visual media.

➠ Models: prithivMLmods/VisionGuardrail-Evo2-27B, prithivMLmods/VisionGuardrail-Evo2-27B-GGUF

➠ Collection: https://huggingface.co/collections/prithivMLmods/visionguardrail-evo2

➠ Previous Models: https://huggingface.co/collections/prithivMLmods/visionguardrail-collection

⤷ To learn more, visit the app page or the respective model pages.
AbstractPhil 
posted an update 17 days ago
view post
Post
45
After a week of failures and invalid hypothesis with bytelex using generic structures, I found a highly successful aleph prototypical structure that conforms to the needs. This structure conforms to standard transformer, FFN, and RNN with some minor tweaks.

You can speak to the model AbstractPhil/alephllm-chat , all of the primary experiments are the listed arms.


AbstractPhil/alephllm-mini-beatrix-training All the weights of the week are stored here and in various nearby directories.

* We've managed to overlap multiple arms to train multiple simultaneous templates.
* Introduce new tokens as composite tokens from multiple teachers.
* Retrain existing tokens into the behavior of one teacher or another.
* Extend new chains and new behaviors from training in combination.
* Properly instantiate and reinforce behavior using Aleph RNN to reinforce training from raw data.

The EMA Relay. The code has been pushed to both beatrix repos.

AbstractPhil/mini-beatrix-2s

The structure itself is built specifically as a solidification unit to extensible arms, allowing more composite structures to build.

EMA structures aren't new, but when applied correctly at just such a methodology, the models begin to behave as though the extension relays are in fact the original model. The chains and behavior form naturally and the substructure begins to conform with the token fragments from much more complex structures like combined token differences of T5, Qwen, and CLIP as unified teachers.

The cross-token noise is mitigated using a series of principles and the blueprints are showing both solidity and failure simultaneously, both proving many new utilizable states and disproving multiple theoretical pathologies utilized in current running modern papers as the methodologies tested in the specific formats.
  • 1 reply
·
prithivMLmods 
posted an update 20 days ago
AbstractPhil 
posted an update 27 days ago
view post
Post
92
The post-beatrix-2s and control variant article is finally satisfactory, so the article is now released https://huggingface.co/blog/AbstractPhil/beatrix-ft2

The control variant will need another train with better SDPA stabilization, as the control variant destabilized and collapsed. The primary fault is the lack of QK normalization, which caused the model to simply collapse given enough time. Claude lists the rest of the suspected reasons in the article.

This was a very difficult series of experiments to tune with many fault points. Trying to make heads or tails of Fable 5.1 Claude-speak hasn't been the easiest task either. It seems the model is more likely to create pedantically rigid responses rather than cooperative. Not necessarily insulting, but definitely a sort of refrigerator-magnet behavior - treating my individual contributions as little sketches for the refrigerator. This often completely ignores my larger MD or complex behavioral instructions in favor of my theoretical or hypothetical - likely considering the MD and technical as the model's own, rather than my direct contributions. Right there... right on the refrigerator goes my hypothesis that worked.

https://github.com/AbstractEyes/geolip-bytelex

In any case, this upcoming week will be related entirely to cross-tokenizer distillation research. It may stretch long beyond the next week, but as it stands the geometric vocabulary has evolved into a codebook prediction system.

I would like to give this program linear wings. The Beatrix model supports it, but how well is up for this week to decide.

There are a multitude of potentials based on a series of very recent articles I will be exploring, providing the necessary bytelex complexity to a roughly 60 hour battery of experiments and trainings throughout the geometric systems.

The results will determine the best and worst methodologies of using these models, these shapes, and these structures with more complex byte-level cross tokenization systems
  • 2 replies
·
prithivMLmods 
posted an update 27 days ago
view post
Post
3847
VisionGuardrail, a multimodal content-safety classifier based on Qwen3.5, is now available on Hugging Face in 4B and 9B variants. It is a direct upgrade to ImageShield-MMCF, providing improved parental controls through conservative visual content-safety filtering.

More About:
➠ hf.co/blog — https://huggingface.co/blog/prithivMLmods/vision-guardrail-mini-blog

➠ Models:
✦ VisionGuardrail-4B: prithivMLmods/VisionGuardrail-4B
✦ VisionGuardrail-9B: prithivMLmods/VisionGuardrail-9B

➠ Dataset:
✦ ImageShield-Guardrail-Pro: prithivMLmods/ImageShield-Guardrail-Pro

⤷ To learn more, visit the app page or the respective model pages.
AbstractPhil 
posted an update about 1 month ago
view post
Post
2039
Mini-Beatrix-2s pretraining is ready.
AbstractPhil/mini-beatrix-2s
The model passed a great deal of rigor and hardship, trained roughly 16 billion tokens or so. The full writeup for the model including the arms training for the first version arms and the second version arms will be drafted and prepared as soon as the v2 arms are done training and testing.

There are many possibilities present with such a model. The hub itself has been marked capable of potentially operating as similarity comparison, 87% of the capacity retained within a 256 dim structure. Along with this, the multi-dimensional hub attention system shows serious promise with controlling diffusion model inference, which I look forward to see the results of.

Additionally, sentence similarity, next token prediction, and a large array of prediction formats have been heavily improved by introducing the full model with splat attention. The model not only improved, the structure complemented everything measured, along with the more effective training regiment for version 2.

Beatrix 2s is essentially an autoregression decoder, however the attention mechanism houses a dual-stage encoder/decoder structure internally. Each adopting the SVAE as a core component, revamped and fitted to the exact rules of AlephLM. So there are essentially 20 SVAE in this structure, each with their own independent encoders, residually learning from the last.

Upcoming tests will include finetunes to bring out the strengths of all special tokens, presented in the upcoming article. The full experiment battery will be completed within a few days and the findings presented.

Modularization, compartmentalization, secularized behavior, and everything between are to be tested with rigor. This model is a rapid learner, there will likely be byproduct problems with that, and I look forward to solving the corewise problems one at a time until the model is strong enough to be useful for all the tested tasks.
  • 4 replies
·
AbstractPhil 
posted an update about 1 month ago
view post
Post
87
Mini-Beatrix-2s is cooking with full splat attention through and through. This model is still trigram, I did this to get a baseline because there's already a trigram model to compare to. This one should be done in a few days and ought to be substantially more intelligent than the first.

Specs are;
Around 220m params, 4096 context window, d1024 model size, splat 128, 1024, 1024, 1024, and so on, 3 experts per block, information banks for storage and retrieval, and a lot of technical knowhow between A to B.

Differences with V2;
Special tokens are implemented byte-directly, so the model will have no problem recognizing an array of special tokens such as DOC, EOF, and a multitude of others.

Suffice it to say, this model is bigger than the first at about 2x. Not just bigger though, estimated to be roughly 8x more intelligent based on the measures.

That being said, the actual model needs to be substantially larger to encompass the full space. The measured space is considerably larger through the small tests for stability, however the full 900m version runs at only around 8k tokens per second with an anchor count of 131,000 and a matching number of heads. This means the full train would require roughly 26 days on a rtx 6000 pro blackwell, which is substantially beyond the expectation curve.

So the smaller one will do for now until I can secure a bit of funding. In any case, the tokenizer system will be implemented on this version after a stable run completes.
  • 2 replies
·
prithivMLmods 
posted an update about 1 month ago
view post
Post
3099
ImageShield-MMCF — Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!

This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.

The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.

⊹ ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
⊹ ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
  • 2 replies
·
prithivMLmods 
posted an update about 1 month ago
view post
Post
5286
The Qwen3.8 27B demo for object grounding is now available on Hugging Face Spaces.

It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).

Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
AbstractPhil 
posted an update about 2 months ago
view post
Post
153
I believe I have a solution for cross-tokenizer chatter and noise, which I've built a prototype repo for this exact tooling dubbed bytelex. https://github.com/AbstractEyes/geolip-bytelex

I had a bit of an inspiration recently and built a prototype for a token translation matrix that I called geolip-bytelex, which allows bytewise translation of many different tokenizers into byte format. The goal is to allow comparative distillation from multiple models to simultaneously represent expertise based on input tokens and differentiated teacher/student InfoNCE and MSE training paradigms, while cutting a huge cost of the distillation analysis comparative compute that cross-tokenizer noise will naturally cause when tokenizers are mismatched or incorrect, reducing a large portion of invalidity from the trained systems established by incorrect valuations from the distillations and lora trainings.

Bytelex is essentially a byte-wise deconstruction of a tokenizer's state into a preliminary 255 byte language allowing for 10s of thousands of sequences per token to be represented rather than just a few. I'm not the first to try this, however I'm in a unique position due to my creation AlephLM being built entirely by learning it's own lexicon, thus allowing this to be more than experiment and instead a working prototype distillation potential.

This can solve a longstanding multi-tokenizer problem that I and many other researchers have been facing, at the cost of setup overhead compute for the preliminary experiments, however the translation matrix I'm planning will potentially solve this problem allowing models to be directly bytewise captured in a more guaranteed methodology through cross-sampled analysis at distillation time in this optimizer state that I'm working out.

I've dubbed this distillation loss ByteInfoNCE and the preliminary is showing humongous promise, with that the bytelex is the crux and prototype concept that I'll be expanding and researching further.
  • 3 replies
·
AbstractPhil 
posted an update about 2 months ago
view post
Post
108
Say hello to the Trigram ByteLLM - AlephLLM: Mini-Beatrix - in her huggingface space!
She is currently stepped at 24000 steps aka 7b tokens in the first couple datasets, so she's not very smart yet.
AbstractPhil/alephllm-chat

Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her.

The AlephLLM prototype is currently in full training with SDPA attention.
AbstractPhil/alephllm-mini-beatrix-training

https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype.
As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults.

It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning.

Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
  • 3 replies
·
AbstractPhil 
posted an update about 2 months ago
view post
Post
679
The upcoming AlephLM LLM prototype "Mini-Beatrix" is based on protocols, rules, and laws established through the process of training AlephLM systems. This will be a first attempt at a smaller full pretrain/finetune of the AlephLLM on raw data, and this will require over a billion unigram tokens.

Mini-Beatrix will inherit an appropriately adapted AlephLM MOE structure containing a multitude of trained experts, a gating system, a long context RoPE system, MHA attention, and a series of hypothesis to answer upon Mini-Beatrix's pretrain and finetune completion.

While focusing on resolving corruptions and invalidity possibly present in the splat attention, the solutions raised SDPA attention protocol token recall ceiling from 0.91 to 0.993. With that the splat attention raised from 0.81 to 0.89~ splat being around 3x the speed is still imperfect.

So far so good. The corruptions have resolved multiple core component overlapping problems causing the AlephLM's inability to handle the trigram system, the structure of the SVAE having faulty trigram structures, and additionally a multitude of other systems in the lineup that were inheriting the corruptions from the core experiment sets.

These corruptions resolved show that the accuracy of standard multiheaded attention will provide the necessary token recall for full LM capacity, and with that If and WHEN I solve the Rorschach Splat attention will be the faster alternative at >=r1 0.99%, only then. The splat attention's considerably larger head count still contains unresolved inconsistencies.

That being said the SDPA MHA attention will be present for the first attempted mini-llm train, which will be named "Mini-Beatrix" with the appropriate sizing associated with this.


The only thing that will change Mini-Beatrix's trajectory will be if Splat attention is perfected between today and next week, which will likely take longer unless I run into a core corruption that has been overlooked through hundreds of analysis.
  • 2 replies
·
AbstractPhil 
posted an update about 2 months ago
view post
Post
2754
The AlephLM results are rolling in and I'm very excited for the possibilities. I am very much looking forward to the coming weeks as I train the first AlephLM distillations from MANY teachers into AMOE arms.

The AMOE arms hook cleanly to AlephLM structures and provide pos/neg learning elements. Hard positive and hard negatives coalesce to extend the capacity.

AbstractPhil/alephlm-0
AbstractPhil/alephlm-adopt-0

As it stands they are structurally sound enough to fully pretrain. As or more stable than a standard Bert experimentally to distill using InfoNCE. AMOE legs improve these structures substantially.

Structural behavior can be expanded in many ways on distilled and pretrained models alike. Attaching the AMOE to any model I've tried has created expanded or improved behavioral accumulations. They do have downsides but their upsides are very experimentally exciting.

I've distilled multiple vits, multiple berts, and have begun distilling berts into AlephLM structures successfully.

This is overall very exciting for me. I've begun formatting larger variants such as including GPT-2 and Qwen 3.5 4b as a paired combinator utilizing pathological T5 learned distilled encodings. It sounds odd, but the results show everything can be expanded and even be taught to cooperate.

The CaptionBert-8192-v2 and v2-b are both structurally collapsing after token 480 or so, which is expected due to the small train. By distilling an AMOE arm to V2 by training with a longformer expert, the results are cutting through like butter. V2 has begun stabilizing rapidly for considerably longer token chains and sequences, the structure is repairing and building reusable capacity.

I have discovered an improved methodology for sampling the AlephLM for text encoder benchmarks, which is predominantly L2 normalized outputs.

Upcoming large paper for the distillation experiments and results within the next week or two. It's going to be a big one.
  • 3 replies
·
prithivMLmods 
posted an update 2 months ago
view post
Post
5568
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.🤗

➠ Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
➠ collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
➠ github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator

⤷ To learn more, visit the app page or the respective model pages.
AbstractPhil 
posted an update 2 months ago
view post
Post
149
AbstractPhil/clip-vitb-mini-distilled
The semi-successful run series on the VIT-B lineup is live and full of useful baseline distillation information for feature + InfoNCE distillation processing as well as direct feature distillation processing, direct InfoNCE distillation, and multiple other tested methods. https://huggingface.co/blog/AbstractPhil/geometric-memory-ft4

This article showcases the baseline utilization and benchmarks of the earlier experiment line's objective and loss structures tested on 12m features for the vit-b baseline. Not the strongest showcase, but the strongest of the champions did show some serious promise.

Next setup will be a directly aligned set based on the loss and objectives decided by the champions in the first runs, for the second run they operate in direct conjunction with the bert-8192 and captionbert-8192 distillation format directly on clip-vit-l features - this time we're including DINOv3 into the mix for it's high potency.

I'm currently extracting 4 clip-vit-l variants for the CC12m features and will be running the next series on the L size, which will give considerably more active and useful features overall within a smaller package.

The captionbert-8192 has a more unique and difficult to tune for pixel processing parity, but I will spend a few days making sure the smaller prototypes fit before I run the large experiments in order to build towards the larger objectives.

Primarily I need to ensure the memory bank aligns correctly and the constellation conforms to the anchors correctly, as this process was not micro managed enough for this run. The results are nonetheless useful and potent.

The process continues until we cover the entire constellation series.
  • 6 replies
·
AbstractPhil 
posted an update 2 months ago
view post
Post
200
The geometric memory article ft4 is live. https://huggingface.co/blog/AbstractPhil/geometric-memory-ft4

Direct pivot to distillation. I've accumulated enough experimental information to directly pivot my long term structure plan to distillation. This is to begin forming entire collectives of cooperative systems; differentiated expert distillation for generative behavior utilizing aleph addressed bottlenecks. With this I've also heavily begun experimenting with aleph competitions and cooperation using multiple pretrained frozen codebooks established from the SVAE system.

The idea here is simple in theory; use InfoNCE and address independent experts to build a manifest of unique gated experts utilizing a multitude of distilled systems from many other models. Such as SigLIP 16B + LAION CLIPB as a pair. The experimentation in the past showed this process is potent and with that merits additional experimentation using the newly established paradigms.

There are quite a bit of experiments to compare these to, so I have no shortage of comparators. After we train our baseline TinyViT with our gated system, we will know which experts are better at what and why they are better.

As a direct continuation from the earlier CLIP distillation experiments I'm directly comparing InfoNCE anchoring with multiple industry standard distillations from multiple papers. First comparison is InfoNCE anchoring in comparison to raw features using CoCo and CLIP_B, which seemed like a fair experiment to train a student with.

The upcoming series of experiments will provide the necessary information for how effective or ineffective this process is.

AbstractPhil/bulk-coco-features

The first experiments will be based on multiple clips from the bulk-coco-features extractions.

First we start with some clips, then some berts, then some smaller qwens, then some larger models, then some much much larger models. All meant to be compacted into selection mechanisms.
  • 4 replies
·