Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
30
P(doom)
Rejects the premise
Follow
Tico1982's profile picture
natalie5's profile picture
senormosca's profile picture
105 followers
·
147 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
replied
to
their
post
about 5 hours ago
Beatrix v3's attention has problems. The model at high byte counts simply cannot recall very well, which has been an expensive lesson but a necessary one. To remedy this fault, I will complete the Beatrix 2s softmax control variant until stable. I will train the curriculum stages with QK normalization similar to Gemma, giving us a proper full model MHA softmax control. The answer is likely faulty bf16, needs fp32 and QK norms on the attention, softmax attention in intervals for the primary model. After stabilizing the v2 softmax variant, I will train the v3 softmax variant as well. These are required. There are many open questions about stability at depth and stability at context window capacity, all of which can be answered with the high accuracy recall at depth and context window size. An expensive lesson, but a necessary lesson for progress. Even with the faults, Beatrix V3's mechanisms have proven to be invaluable in their own rite. Multiple experiments have yielded since, including the tokenizer processing system - giving Beatrix the ability to directly speak as another model's tokenizer language. It's not perfect yet, but the process is likely reusable for other similar architected models. The quiet loss has been extensively successful. The structure of each arm trained one after the other shows we can keep each arm nice and quiet while being fully expert in their own field. Activation happens when the byte pattern says so, and the arms stay nice and quiet when the gated pattern doesn't say so. Modularized arms are successful. This is a process that can likely be utilized on other models as well, but the process hasn't been fully completed yet. The model Beatrix was a good catalyst for this because of the byte format of the model and the anchored behavior allows this to be trained more quickly and stable. In any case, the 2s-control is upcoming and will begin today. I'll try to get it cleaned up by the end of the week to prepare for the 3s control.
updated
a model
about 7 hours ago
AbstractPhil/alephllm-mini-beatrix-training
replied
to
their
post
about 10 hours ago
Beatrix v3's attention has problems. The model at high byte counts simply cannot recall very well, which has been an expensive lesson but a necessary one. To remedy this fault, I will complete the Beatrix 2s softmax control variant until stable. I will train the curriculum stages with QK normalization similar to Gemma, giving us a proper full model MHA softmax control. The answer is likely faulty bf16, needs fp32 and QK norms on the attention, softmax attention in intervals for the primary model. After stabilizing the v2 softmax variant, I will train the v3 softmax variant as well. These are required. There are many open questions about stability at depth and stability at context window capacity, all of which can be answered with the high accuracy recall at depth and context window size. An expensive lesson, but a necessary lesson for progress. Even with the faults, Beatrix V3's mechanisms have proven to be invaluable in their own rite. Multiple experiments have yielded since, including the tokenizer processing system - giving Beatrix the ability to directly speak as another model's tokenizer language. It's not perfect yet, but the process is likely reusable for other similar architected models. The quiet loss has been extensively successful. The structure of each arm trained one after the other shows we can keep each arm nice and quiet while being fully expert in their own field. Activation happens when the byte pattern says so, and the arms stay nice and quiet when the gated pattern doesn't say so. Modularized arms are successful. This is a process that can likely be utilized on other models as well, but the process hasn't been fully completed yet. The model Beatrix was a good catalyst for this because of the byte format of the model and the anchored behavior allows this to be trained more quickly and stable. In any case, the 2s-control is upcoming and will begin today. I'll try to get it cleaned up by the end of the week to prepare for the 3s control.
View all activity
Organizations
AbstractPhil
's models
223
Sort: Recently updated
AbstractPhil/alephllm-mini-beatrix-training
Updated
3 minutes ago
•
1
AbstractPhil/beatrix-tokenizers
Updated
about 11 hours ago
AbstractPhil/mega-liminal-lora
Text-to-Image
•
Updated
1 day ago
AbstractPhil/mini-beatrix-1
Text Generation
•
0.1B
•
Updated
1 day ago
•
1.66k
•
1
AbstractPhil/mini-beatrix-2s
Text Generation
•
0.2B
•
Updated
1 day ago
•
1.65k
AbstractPhil/mini-beatrix-2.5s
Text Generation
•
0.2B
•
Updated
1 day ago
•
303
AbstractPhil/mini-beatrix-3
Text Generation
•
0.4B
•
Updated
1 day ago
•
1.33k
AbstractPhil/clip-vitb-mini-distilled
Image Feature Extraction
•
8.93M
•
Updated
1 day ago
•
64
AbstractPhil/geolip-beatrix-anima
Text-to-Image
•
Updated
1 day ago
•
1
AbstractPhil/geolip-beatrix-sana
Text-to-Image
•
Updated
6 days ago
AbstractPhil/sd15-flow-lune
Text-to-Image
•
Updated
11 days ago
•
53
AbstractPhil/geolip-bytelex
Updated
24 days ago
AbstractPhil/aleph-splat-0
Updated
Sep 5
AbstractPhil/rnn-cifar10-t0
Updated
Aug 28
AbstractPhil/alephlm-0
Feature Extraction
•
Updated
Aug 26
AbstractPhil/alephlm-adopt-0
Text Generation
•
Updated
Aug 8
AbstractPhil/captionbert-8192-v2-b
Feature Extraction
•
58.3M
•
Updated
Aug 8
•
25
•
2
AbstractPhil/captionbert-8192-v2
Feature Extraction
•
58.3M
•
Updated
Aug 8
•
36
•
1
AbstractPhil/loss-manifest
Updated
Aug 2
AbstractPhil/geolip-bertenstein
Feature Extraction
•
Updated
Jul 31
AbstractPhil/geolip-vit-captionbank-coco
Image Feature Extraction
•
Updated
Jul 28
AbstractPhil/geolip-vit-base-x3
11.7M
•
Updated
Jul 28
•
12
AbstractPhil/geolip-vit-large-x3
78.3M
•
Updated
Jul 28
•
9
AbstractPhil/geolip-aleph-diffusion
Updated
Jul 25
•
2
AbstractPhil/geolip-aleph-qwen-3.5-0.8b-instruct
Updated
Jul 25
•
1
AbstractPhil/amoe-lora
Updated
Jul 21
AbstractPhil/aleph-diffusion-adapters
Updated
Jul 19
AbstractPhil/qwen3.5-0.8b-relay-caption
Updated
Jul 18
AbstractPhil/geolip-aleph-qwen
Updated
Jul 15
AbstractPhil/geolip-aleph-differentiation
Updated
Jul 12
Previous
1
2
3
...
8
Next