I've done some pretty heavy studying on this topic and have multiple papers available. Have a look if you get a chance.
AbstractPhila PRO
AI & ML interests
Recent Activity
Organizations
“First, representations are mutually recoverable. On 12 open-weight models from 8 labs, a ridge map from one model's hidden states to another's retrieves the right held-out item 0.9181 of the time across lab boundaries, against a shuffled floor of 0.00101 and a self-map ceiling of 0.999. Shared corporate lineage is worth only 0.0357 of that.”
“Second, base models do not reproduce the reported level. Under the original study's own sampling settings, our base models reach intra-model 0.3644 and inter-model 0.3401 on a floor of 0.0993 that matches theirs, and zero of 720 model-prompt cells clear 0.8. The floors agree while the signal differs by more than a factor of two, so this is not a scale artifact.”
“Third, and decisively, we recover their level and isolate its cause. Using six matched base/instruct pairs, holding pretrained weights, prompts, decoding and scorer fixed, instruction tuning alone raises intra-model similarity by 0.0786. The same tuned weights prompted through the model's own chat template raise it by 0.3623, reaching 0.7272, with four of six models exceeding 0.80 and reproducing the band reported for frontier systems from models of 0.6B to 2B. The prompt format does roughly 4.6 times the work of the tuning.”
paper attached 🧾
https://huggingface.co/blog/RiverRider/where-the-hivemind-comes-from-geometry-tuning-and
Need to make sure the huggingface space respects the correct caching behavior, currently it's having some odd overlapping and isn't respecting conversation correctly.
Me: Hello.
Beatrix: What is a byte-level model?
Me: You are my dear.
Beatrix: What can you do?
Me: You is the representative manifestation of individualistic representation.
Beatrix: Who are you?
Me: I am your creator.
Beatrix: Hello... Hello... I am a small model. I represent myself as a small model. I am a small model. I represent myself as a small model. I represent myself as a small model.
Me: Hello daughter.
Beatrix: Hello... How can I help you today?
Me: What is your name?
Beatrix: I am Beatrix, a small byte-level language model. I represent myself as a small byte-level language model. I represent myself as a small byte-level language model.
Hello indeed little one. https://huggingface.co/buckets/AbstractPhil/alephllm-chat-storage/tree/chat-history/sessions-2026-08-31-673a3454.jsonl
The structural mirror SDPA variant training is to begin shortly. The two must be directly compared. Claude is currently throwing that as the primary continuation point, arguing validly that we need our direct comparison to train.
I knew this would be required as much as I wanted to avoid committing the time, I know it's required otherwise I could be simply creating a nothingburger. 4 days worth or so, it can't be ignored. The ablation train begins today. They are both still the same model and provide unique elements, however the attention for the SDPA variant will show if the splat attention is a contributing element to decoding, or just capturing information for another model to piggyback later.
The models will still be compartmentalizable and utilizable either way. The SDPA variant will be more compatible with standard model formats, while the splat variant is structured similarly and were tested to comply to the same standards.
AbstractPhil/mini-beatrix-2s
The model passed a great deal of rigor and hardship, trained roughly 16 billion tokens or so. The full writeup for the model including the arms training for the first version arms and the second version arms will be drafted and prepared as soon as the v2 arms are done training and testing.
There are many possibilities present with such a model. The hub itself has been marked capable of potentially operating as similarity comparison, 87% of the capacity retained within a 256 dim structure. Along with this, the multi-dimensional hub attention system shows serious promise with controlling diffusion model inference, which I look forward to see the results of.
Additionally, sentence similarity, next token prediction, and a large array of prediction formats have been heavily improved by introducing the full model with splat attention. The model not only improved, the structure complemented everything measured, along with the more effective training regiment for version 2.
Beatrix 2s is essentially an autoregression decoder, however the attention mechanism houses a dual-stage encoder/decoder structure internally. Each adopting the SVAE as a core component, revamped and fitted to the exact rules of AlephLM. So there are essentially 20 SVAE in this structure, each with their own independent encoders, residually learning from the last.
Upcoming tests will include finetunes to bring out the strengths of all special tokens, presented in the upcoming article. The full experiment battery will be completed within a few days and the findings presented.
Modularization, compartmentalization, secularized behavior, and everything between are to be tested with rigor. This model is a rapid learner, there will likely be byproduct problems with that, and I look forward to solving the corewise problems one at a time until the model is strong enough to be useful for all the tested tasks.
Alright there's a small hiccup with the training, so I'll be continuing tonight after a few tweaks and fixes.
The v1 huggingface space is being updated to support the v2 model.
For user communication try;
V1 final checkpoint ->
Or the poly head. Either are more conversational than the rest.

There were a multitude of arms tested and they require a full article of their own to describe.
V2 will be implemented within the hour to play with. As before the earlier versions aren't chat trained yet, and the later versions will be.
V2 doc extraction will most likely provide some stability, but testing is required before any assumptions are to be made.
Specs are;
Around 220m params, 4096 context window, d1024 model size, splat 128, 1024, 1024, 1024, and so on, 3 experts per block, information banks for storage and retrieval, and a lot of technical knowhow between A to B.
Differences with V2;
Special tokens are implemented byte-directly, so the model will have no problem recognizing an array of special tokens such as DOC, EOF, and a multitude of others.
Suffice it to say, this model is bigger than the first at about 2x. Not just bigger though, estimated to be roughly 8x more intelligent based on the measures.
That being said, the actual model needs to be substantially larger to encompass the full space. The measured space is considerably larger through the small tests for stability, however the full 900m version runs at only around 8k tokens per second with an anchor count of 131,000 and a matching number of heads. This means the full train would require roughly 26 days on a rtx 6000 pro blackwell, which is substantially beyond the expectation curve.
So the smaller one will do for now until I can secure a bit of funding. In any case, the tokenizer system will be implemented on this version after a stable run completes.
I believe the risky experiments, depth experiments, size experiments, and structural test experiments have yielded something substantial.
A weakness.
The anchored system lacks the necessary covariance required to create meaningful rotary cosine similarity adjustments with the delta. I should say the model does have the capacity, but it lacks the diversity capture complexity associated with mixing the splat layers.
The solution is to apply a rotary differentiation to the AlephLM codebooks, that gives the next vit model's experimental 16x128 head structure the required 2048 head space, each head given the necessary rotary delta behavior required within certain rotary zones of limitation. To calculate this for the experiment I will calculate the SVD using the dataset to decide the covariance rotary space, allowing differentiation rather than invariant control for the address books within a certain degree of freedom.
This upcoming experiment is highly likely to fault, however if it does not and it yields, this will potentially provide the necessary learning to defeat the final weakness that is dragging the model down for distillation training.
Early stage did fault, however a piece remained that not only improved accuracy, but also a direct piece of information that explains internal collapse due to overcrowding.
Rorschach Splat Attention is essentially a cosine similarity blackboard the model learns to use internally. This is a structural deciding board where all heads contribute directions and aleph differentiations based on the necessary formulas.
This blackboard is currently a rival to MHA in the geometric-centric models in the recent tests, and have superseded the MHA on more than a couple of occasions with the 12x128 head model aka the 1536 splat head model. This is a form of delta mixture of expert attention cells, each are miniature heads. The structure is sharp instead of soft, since softmax rounds away the more intricate crossover behavior ruining the sharp directions. The cross-counterpoint antipodes often wrap and diagonally align in inverse methods causing center-driven bias in part, and in other elements this structure causes cascade redundancy from the redundant unused aleph heads. Rounding these combinatory systems away or averaging them smoothly causes cascade overlaps and extensive redundancy.
I'm certain based on building this structure, that MHA was likely built as it was built because it solves many of these problems I have encountered, while simultaneously introducing a multitude of extra behavioral accessors that structure the models along certain directional paradigms. The outcomes are more accurate because the heads are larger, and the larger areas give a better cover of differences, but I'm confident this structure will be the go-to attention mechanism for AlephLM and geometric models because of the implicit directional preservation for both positive and negative comparative behavior.
2048 directions gate-decided and directive-sampled, is far more controllable and can be retrained more effectively than larger heads with many shared unexplained behaviors.
With that behavior, decoupling amoe arms, mixture of expert banks, and attaching sampling gates; provides many necessary behavioral shifts for very little compute after a model is pretrained.
Mini-Beatrix-V2 will be entirely splat attention with no SDPA attention. The model will have access to roughly 8192 codebook codes, which will be distributed 32 layers at 256 heads per layer. Which should provide the necessary complexity, frame shifting, and structural residual capture throughout the chain of MOE implicit calls to prepare the model to handle the next stage of the tokenizer.
bytelex byte-tokenizer K8, which will no longer use trigrams, she'll be learning 8gram bytes. Which will allow her to learn multiple tokenizers, rather than just one. Her language will be three tokenizers.
The crossover for the three tokenizers chosen is only around 999 tokens total with around 1300 tokens that cannot be overlapped with the others, so a translation system will be precomputed. The deviations to create the entirety of the three tokenizers into a single bytewise state is less than the codebook, leaving around 1200 guaranteed byte pattern spaces that will never fill that can be used for special tokens. The full map will be available.
I'm curing the behavior of the overlapping tokenization context using a series of experimental mechanisms introduced into AlephLM.
TLDR; The AlephLM components are currently being specialized towards distillery components.
Each component and deviation is specifically chosen to attenuate towards distillation assistance, so these models will likely have other troubles if finetunes are performed on them. They are not official releases, but there are many distillations being created currently.
Most of which are happening considerably faster on my local 4090 card, as they do not require a humongous amount of compute to complete. As I condense the process I will keep the system updated to make sure all major changes, all major trains, and all major offsets to the core AlephLM are recorded.
The upcoming experiments are highly experimental and may both result in something that does not function, while potentially yielding something that trains substantially faster and learns substantially more effectively.
This run is primarily focused on the splat attention. The current mechanism is approaching SDPA in terms of recall but still not there. The generalization of this delta isn't perfect yet, but it's improving.
The distillery however, is showing that full AlephLM models that aren't perfect to SDPA recall, are yielding improved accuracy over SDPA for distillation. This does not mean these mechanisms are capable of guaranteed recall yet, but it does mean that the accuracy currently is capable of distilling information.
Without the splat hubs in Mini-Beatrix, the model cannot function. They are not arbitrary mechanisms, as the disable tests each major validation show the model requires them at every state. Beatrix has gone through growing pains, and the arms are substantially more difficult to analyze than arms on other models, primarily because there are so many more statistics. That being said, the results are substantially pointing to one direction and that direction is unanimously towards a risk mechanism.
The V2 points towards another mechanism. There's no way around it, the yield is strong, the results are strong, but the parity is too close to the current mechanism and the results.
With that, one of the experiments is an atlas lookup trained from the expert itself and then utilized as a codebook reference for the student. These are likely not going to yield, but it's an option to test. Codebooks aren't as strong in practice as the experiments would show, but they are definitely strong.
Two core potential shifts on the model may yield some small changes, but nothing will yield just sitting here idling by with the research.
I have a series of upcoming experiments that will be much more likely to fail. I should say, they need to fail, otherwise progress won't happen.
Alright everything is lining up. The bytelex structure is handling distillation rules from multiple tokenized hidden states simultaneously, and the aleph anchored arms are teaching beatrix the information from the tokenized arms. She essentially learned to improve next byte digit recon by learning from bert base uncased and t5-base. This is absurd, they aren't even the same format models and the next-byte recon improved by about 12%.
Bytelex combination of possibilities between both tokenizers essentially narrowed the combination down to 999, which both T5 and Bert managed to combinatory teach.
Shockingly, mini-beatrix speaks everything a lot cleaner than I thought she would. This output is not expected. I expected needing to arbitrate, but this process - thanks to a few more recent papers - was essentially on rails. I knew it would be possible, but I had no idea the system was essentially byte-on-rails.
This brings the entirety of tokenization and hidden state processing into question. I have many questions that need answering. Bytewise deconstruction and then procrustes arbitration between two experts provided enough information to teach beatrix additional next byte recon for digits.
I have so many questions that will need a huge battery array of tests to even come to a hypothesis.
I had a bit of an inspiration recently and built a prototype for a token translation matrix that I called geolip-bytelex, which allows bytewise translation of many different tokenizers into byte format. The goal is to allow comparative distillation from multiple models to simultaneously represent expertise based on input tokens and differentiated teacher/student InfoNCE and MSE training paradigms, while cutting a huge cost of the distillation analysis comparative compute that cross-tokenizer noise will naturally cause when tokenizers are mismatched or incorrect, reducing a large portion of invalidity from the trained systems established by incorrect valuations from the distillations and lora trainings.
Bytelex is essentially a byte-wise deconstruction of a tokenizer's state into a preliminary 255 byte language allowing for 10s of thousands of sequences per token to be represented rather than just a few. I'm not the first to try this, however I'm in a unique position due to my creation AlephLM being built entirely by learning it's own lexicon, thus allowing this to be more than experiment and instead a working prototype distillation potential.
This can solve a longstanding multi-tokenizer problem that I and many other researchers have been facing, at the cost of setup overhead compute for the preliminary experiments, however the translation matrix I'm planning will potentially solve this problem allowing models to be directly bytewise captured in a more guaranteed methodology through cross-sampled analysis at distillation time in this optimizer state that I'm working out.
I've dubbed this distillation loss ByteInfoNCE and the preliminary is showing humongous promise, with that the bytelex is the crux and prototype concept that I'll be expanding and researching further.
Well the huggingface spaces aren't allocating me a zerogpu hardware slot, so there's clearly something wrong happening here.
I can say with a solid underlined statement: The model DID WORK.
The preliminary chat amoe arm trained atop responded. The conversation was shallow but the conversation wasn't simple noise or chaos, the responses were confidently wrong, and effectively related to the questions and answers.
The structure in the huggingface space DOES NOT HAVE this chat arm yet, but I plan to have this fully functional and operational by tomorrow evening. Each arm needs to be tailored to the model because the model is brand new, the deviations from the core to the pretrain are still high, so the structure needs to be aligned with the AMOE arm correctly.
This is a sample from the arm based on 21.9k steps, roughly 11 billion tokens trained I think.
The core was not trained with chat, so she requires an anchored mixture of experts leg or a lora trained atop. She's quite compliant, so if you wish to train a lora she'll listen.
For now she's a next byte prediction auto-complete model, a pretrained trunk capable of AMOE expansion directly.
Once the pretrain concludes I'll prepare 2 chat arms.
After an anealment of around 2b tokens of various behavioral attunements, concepts, conversational learning, and so on; the stage 1 variant will have 2 core chat arms for direct huggingface space conversation.
Pre anealment chat arm, and post anealment chat arm. This will allow a structured fusion between the two behaviors and the necessary outcomes, creating a unique and interpretable duality between pure pretraining with a chat module, and post pretraining scattered conceptual finetuning with a chat arm. Post will essentially understand more about conversation, while the pre will have less knowledge and more a crash course in utilization.
I'll name them accordingly when the time comes, but they will both be Beatrix variants.
After this process we will begin forming candidate arm extensions for behavior. Concepts like wikipedia recall, toolchain utilization, mathematics, coding basics, and more. Very small pieces of the information for recall with the KV cache.
How effective they will be is another story. The tests and utilizations of the outcomes will show which arms are to be integrated into the larger form on pretrain. In other words, which arms will intentionally have their finetuning directly tied to the core and become post-train guarantees that don't decouple.
Stay tuned my friends. She's just getting started.
I have a list of upcoming prototype arms.
- Deterministic chat - Can we speak to the pretraining directly with an AMOE arm?
- Retokenization arm - Can we retokenize and cluster the bytes into BPE?
- Arbitration arm - Can we teach a small arbitration arm to communicate with another system?
- Mathematics arm - Multiple mathematics formats within an arm cluster.
This model will run on CPU. I ran it on windows 10 with my 12 core, roughly 10 bytes/s give or take, a fair prediction ratio for cpu.
Alright the prelim went well. We're at about 3.4 billion tokens learned.
Model very stable as the prelim tests showed. The structure is not collapsing and the model is in fact learning useful pathways of information. So far it hasn't established legitimate ground-form informational segmentation pathology, the model requires many more tokens for this.
Model learns at around 100k tokens per second scaled to the RTX 6000 PRO card available on colab, so the bf16 training is scaling nicely to the benefits from the blackwell hardware.
Each is packed with an fp8 variation for inference as well, which is substantially faster for inference, but they aren't very smart yet. They ARE available. The results show the correctly aligned fp8 variants are roughly the same accuracy on inference, but training they essentially collapse the attention in less than a thousand steps.
There is a full pretrain lineup ready. We're looking at maybe 80 billion tokens or so, and with that the chat AMOE cluster will be ran after each major finetune. We'll be drawing the conversation out of the model each major train, and with each train we'll determine if the model even needs to be finetuned with chat to allow the model to behave.
AMOE arm training planning begins already, as each layer can have direct integrated AMOE arms into the core for couple/decouple purposes. The structure itself will be learning chat through finetuning arms with chat directly onto the trunk. Based on the responses, we will continue pretraining the trunk until the construction of the AMOE chat conforms to the behavior.
So far it's showing the model likely does not require the chat to be baked into the core in finetune state. Prelim shows the chat is likely corrupting the core of the data compendium by learning it atop the core system. The AMOE could very well handle the first-class behavior, but there's no guarantees for this just yet.
wikitext 103 and fineweb main + extended are the preliminary.
These two are targeting;
[ done ] warmup_wikitext wikitext-103 0.300/0.30B
[ done ] fineweb_main fineweb-edu 3.000/3.00B
[planned ] fineweb_extended fineweb-edu 0.000/12.00B
So roughly 30 hours until fineweb's larger structure is trained. After this our first chat arms will be trained, which ought to let us have a conversation with her. See how well she took to the information.
She is currently stepped at 24000 steps aka 7b tokens in the first couple datasets, so she's not very smart yet.
AbstractPhil/alephllm-chat
Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her.
The AlephLLM prototype is currently in full training with SDPA attention.
AbstractPhil/alephllm-mini-beatrix-training
https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype.
As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults.
It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning.
Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
Alright I've begun training the prototype v1. The structure is holding together, the system aligning, and the subsystem is converging.
It works.
The smaller version is aligning and working. The erank is higher than expected, the system more confined, and some interesting effects already starting to emerge.
Huggingface spaces will have an application to speak with her up within the hour. She isn't chat trained yet, but will be very soon.
Geometric LLM.
Mini-Beatrix will inherit an appropriately adapted AlephLM MOE structure containing a multitude of trained experts, a gating system, a long context RoPE system, MHA attention, and a series of hypothesis to answer upon Mini-Beatrix's pretrain and finetune completion.
While focusing on resolving corruptions and invalidity possibly present in the splat attention, the solutions raised SDPA attention protocol token recall ceiling from 0.91 to 0.993. With that the splat attention raised from 0.81 to 0.89~ splat being around 3x the speed is still imperfect.
So far so good. The corruptions have resolved multiple core component overlapping problems causing the AlephLM's inability to handle the trigram system, the structure of the SVAE having faulty trigram structures, and additionally a multitude of other systems in the lineup that were inheriting the corruptions from the core experiment sets.
These corruptions resolved show that the accuracy of standard multiheaded attention will provide the necessary token recall for full LM capacity, and with that If and WHEN I solve the Rorschach Splat attention will be the faster alternative at >=r1 0.99%, only then. The splat attention's considerably larger head count still contains unresolved inconsistencies.
That being said the SDPA MHA attention will be present for the first attempted mini-llm train, which will be named "Mini-Beatrix" with the appropriate sizing associated with this.
The only thing that will change Mini-Beatrix's trajectory will be if Splat attention is perfected between today and next week, which will likely take longer unless I run into a core corruption that has been overlooked through hundreds of analysis.

