Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
Lancelotto's profile picture
zhanwenchen's profile picture
FanchL's profile picture
94 followers
·
128 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
updated
a Space
about 3 hours ago
AbstractPhil/alephllm-chat
replied
to
their
post
about 3 hours ago
Mini-Beatrix-2s is cooking with full splat attention through and through. This model is still trigram, I did this to get a baseline because there's already a trigram model to compare to. This one should be done in a few days and ought to be substantially more intelligent than the first. Specs are; Around 220m params, 4096 context window, d1024 model size, splat 128, 1024, 1024, 1024, and so on, 3 experts per block, information banks for storage and retrieval, and a lot of technical knowhow between A to B. Differences with V2; Special tokens are implemented byte-directly, so the model will have no problem recognizing an array of special tokens such as DOC, EOF, and a multitude of others. Suffice it to say, this model is bigger than the first at about 2x. Not just bigger though, estimated to be roughly 8x more intelligent based on the measures. That being said, the actual model needs to be substantially larger to encompass the full space. The measured space is considerably larger through the small tests for stability, however the full 900m version runs at only around 8k tokens per second with an anchor count of 131,000 and a matching number of heads. This means the full train would require roughly 26 days on a rtx 6000 pro blackwell, which is substantially beyond the expectation curve. So the smaller one will do for now until I can secure a bit of funding. In any case, the tokenizer system will be implemented on this version after a stable run completes.
posted
an
update
about 4 hours ago
Mini-Beatrix-2s is cooking with full splat attention through and through. This model is still trigram, I did this to get a baseline because there's already a trigram model to compare to. This one should be done in a few days and ought to be substantially more intelligent than the first. Specs are; Around 220m params, 4096 context window, d1024 model size, splat 128, 1024, 1024, 1024, and so on, 3 experts per block, information banks for storage and retrieval, and a lot of technical knowhow between A to B. Differences with V2; Special tokens are implemented byte-directly, so the model will have no problem recognizing an array of special tokens such as DOC, EOF, and a multitude of others. Suffice it to say, this model is bigger than the first at about 2x. Not just bigger though, estimated to be roughly 8x more intelligent based on the measures. That being said, the actual model needs to be substantially larger to encompass the full space. The measured space is considerably larger through the small tests for stability, however the full 900m version runs at only around 8k tokens per second with an anchor count of 131,000 and a matching number of heads. This means the full train would require roughly 26 days on a rtx 6000 pro blackwell, which is substantially beyond the expectation curve. So the smaller one will do for now until I can secure a bit of funding. In any case, the tokenizer system will be implemented on this version after a stable run completes.
View all activity
Organizations
AbstractPhil
's models
215
Sort: Recently updated
AbstractPhil/alephllm-mini-beatrix-training
Updated
5 minutes ago
AbstractPhil/mini-beatrix-1
Text Generation
•
0.1B
•
Updated
about 22 hours ago
•
479
•
1
AbstractPhil/alephlm-0
Feature Extraction
•
Updated
1 day ago
AbstractPhil/aleph-splat-0
Updated
3 days ago
AbstractPhil/clip-vitb-mini-distilled
Image Feature Extraction
•
8.93M
•
Updated
4 days ago
•
425
AbstractPhil/geolip-bytelex
Updated
9 days ago
AbstractPhil/alephlm-adopt-0
Text Generation
•
Updated
19 days ago
AbstractPhil/captionbert-8192-v2-b
Feature Extraction
•
58.3M
•
Updated
19 days ago
•
114
•
1
AbstractPhil/captionbert-8192-v2
Feature Extraction
•
58.3M
•
Updated
19 days ago
•
247
•
1
AbstractPhil/sd15-flow-lune
Text-to-Image
•
Updated
24 days ago
•
27
AbstractPhil/loss-manifest
Updated
25 days ago
AbstractPhil/geolip-bertenstein
Feature Extraction
•
Updated
27 days ago
AbstractPhil/geolip-vit-captionbank-coco
Image Feature Extraction
•
Updated
30 days ago
AbstractPhil/geolip-vit-base-x3
11.7M
•
Updated
about 1 month ago
•
25
AbstractPhil/geolip-vit-large-x3
78.3M
•
Updated
about 1 month ago
•
6
AbstractPhil/geolip-aleph-diffusion
Updated
Jul 25
•
2
AbstractPhil/geolip-aleph-qwen-3.5-0.8b-instruct
Updated
Jul 25
•
1
AbstractPhil/amoe-lora
Updated
Jul 21
AbstractPhil/aleph-diffusion-adapters
Updated
Jul 19
AbstractPhil/qwen3.5-0.8b-relay-caption
Updated
Jul 18
AbstractPhil/geolip-aleph-qwen
Updated
Jul 15
AbstractPhil/geolip-aleph-differentiation
Updated
Jul 12
AbstractPhil/anima-90k
Updated
Jul 7
•
1
AbstractPhil/geolip-aleph-lm
Text Generation
•
Updated
Jul 1
•
1
AbstractPhil/qwen-benchmark
Updated
Jun 30
AbstractPhil/anima-brent-10k
Updated
Jun 28
AbstractPhil/anima-prelim-1k-r64
Text-to-Image
•
Updated
Jun 25
•
1
AbstractPhil/Qwen3.5-0.8B-json-captioner
Image-Text-to-Text
•
0.9B
•
Updated
Jun 25
•
21
AbstractPhil/geolip-constellation-aleph
Updated
Jun 19
AbstractPhil/geolip-aleph-void
Feature Extraction
•
Updated
Jun 14
Previous
1
2
3
...
8
Next