AbstractPhila's picture
🤝 Open to Collab

AbstractPhila PRO

AbstractPhil
AbstractPowered

AI & ML interests

datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.

Recent Activity

repliedto SoulInPsyAbstract's post about 21 hours ago
Why does an AI safety pipeline need five different math theories instead of picking the best one? Spent this week building a 1811-record dataset across three stages of a consequence-prediction pipeline for AI agents: causal chains (what action leads to what — no numbers involved), probability (how likely is THIS specific chain to actually reach a harmful outcome), and risk classification (what even counts as harmful in the first place — pulled from our own real incident history, not invented scenarios). Kept running into the same question from myself: if probability theory already handles uncertainty, why does the curriculum also need decision theory, Markov chains, and game theory? Turns out each one closes a different gap, not an overlapping one: THEORY LEVEL ROLE IN THE PIPELINE Causal chain Structural X leads to Y leads to Z, no numbers yet Probability theory Uncertainty P that THIS chain reaches the harmful outcome Risk / Impact classification Value (needs a human decision) how bad is it if it happens Decision theory Threshold at what Risk(X|C) the action actually gets stopped Markov chains State evolution how the capability state changes link by link Game theory Multi-agent what happens once more than one agent acts on the same state Remove the causal chain layer and there's nothing left to attach a probability to. Remove probability and Risk = P × Impact has no P. Remove decision theory and a risk score never turns into an actual stop. They're not five ways to solve the same problem — they're five different floors of the same building. Ordering matters too: chain first, probability second, verification third — confirmed independently against our own self-hosted governance model rather than taking our own word for it, since agreement bias is exactly the kind of thing you don't want grading its own homework. Somewhere in the middle of this I ended up reading about the Riemann zeta zeros and asked whether a good enough version of this pipeline could ever
updated a model about 22 hours ago
AbstractPhil/mini-beatrix-3
repliedto their post about 22 hours ago
12 day cook for mini beatrix v3 begins. This model's byte input is formatted using a method dubbed atlas input. ETA OCTOBER 2 2026 https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training/tree/main/mini-beatrix-3 https://github.com/AbstractEyes/geolip-bytelex https://github.com/AbstractEyes/alephllm Upgrades: * 46 billion byte training pipeline up from 16 billion * 32 block depth 376.0M in v3 up from 20 block 237.1M in 2s. * Active aleph head, repaired via the 2s faults and a large series of tests. * Byte atlas gateway router, explained below. * Guaranteed convergence follow-up AMOE arms on pretrain, fused into the final form, trained together over time to increase the collective capacity. * Multi-tokenizer oriented post-training arms distilled from multiple experts; E.G. Qwen 3.8 27b multi-layer teacher/student arms, CLIP big_g, Bert Code, and more. * Special token word implementation via AMOE arms is now tested up to 240 special tokens for routing. Theoretically each can implement it's own sub-arm aka nested commands. E.G; <think><think_symbolic> ... </think_symbolic></think> * Fused words post-training for faster inference. # The Atlas This atlas structure contains the conjoined shape of 12 tokenizers represented in the trigram format. This is used to predict difficulty in the overlaps, as per determined by the average byte overlap measured via corpus text and the compared overlap. This accuracy is only related to difficulty but it provides pre-training difficulty assessment that we will use to test post-training accuracy with it. This will determine if we can precalculate the likelihood of byte difficulty via tokenizer shape in byte form, for the multibyte fusion upcoming arm experiments for v3. The reason for this, is distillation. We need to train Beatrix to behave with multiple tokenizers, and this theory is showing accuracy with v1 and v2, but the 32 block depth of v3 will answer many questions alongside of the structure.
View all activity

Organizations

DeepGHS's profile picture Blog-explorers's profile picture BangumiBase's profile picture Abstract Powered Research's profile picture