Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
13
Qozimo
19
3
Follow
maeuryhoisj's profile picture
peterleeishere's profile picture
2 followers
·
1 following
AI & ML interests
None yet
Recent Activity
updated
a collection
1 day ago
models
replied
to
Undi95
's
post
4 days ago
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek. I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use? No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note. The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match. Later, code will be checked by actually running tests. The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights. Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again. The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move. I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks. Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try. Did you already tried something like that? What was your result? I'm curious!
replied
to
Undi95
's
post
4 days ago
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek. I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use? No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note. The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match. Later, code will be checked by actually running tests. The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights. Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again. The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move. I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks. Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try. Did you already tried something like that? What was your result? I'm curious!
View all activity
Organizations
None yet
Qozimo
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
deepseek-ai/DeepSeek-V4.1-Flash
11 days ago
"Le Cerveau dans un Bocal de Morve: DeepSeek-V4.1-Flash or the Art of Selling a 4B Invalid as a Frontier Thinker" 🧠💧🧪
❤️
4
5
#36 opened 30 days ago by
Qozimo
New activity in
Altworld/Hemmingway-1
17 days ago
How to turn a top-tier model into an intellectual invalid
👀
🤯
4
4
#7 opened 17 days ago by
Qozimo
New activity in
Qwen/Qwen-Image-2.1-PE-I2I
20 days ago
Stop the fake naming hype: Qwen-Image-2.1-PE is just Qwen3.5-9B with stripped MTP layers and a lazy system prompt
🤯
😔
2
5
#1 opened 20 days ago by
Qozimo
New activity in
Qwen/Qwen-Image-2.1
20 days ago
stop doing garbages!
6
#13 opened 20 days ago by
Qozimo
New activity in
insraq/Qwen3.5-2B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated
25 days ago
Garbage!
2
#1 opened 25 days ago by
Qozimo
New activity in
Qwen/Qwen3.8-Flash-Next
28 days ago
"La Grande Illusion de 360 Go: A 4B Dummy in a 512-Expert Costume, Walking on a Stolen N-Gram Cane"
🤯
😔
8
20
#24 opened about 2 months ago by
AdrienneNoctis
New activity in
Qwen/Qwen3.8-27B
about 1 month ago
Le Cirque du Raisonnement: 27B Parameters, 262K Tokens of Hot Air, and a Packaging Bug That Silently Truncates Every Prompt)
❤️
➕
5
19
#144 opened about 2 months ago by
AdrienneNoctis
New activity in
AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
about 1 month ago
Feature Request / Architecture Synthesis for Qwen3.8 27B GA Cut: Validating Low-Distortion Rank-1 Abliteration under Real-World High-Entropy Inference
1
#5 opened about 1 month ago by
Qozimo
New activity in
zai-org/GLM-5.3-Flash
about 1 month ago
A Fresh Serving of "Amazing" Direct from the Shovel: The GLM-5.3-Flash Illusion
1
#28 opened about 1 month ago by
Qozimo
New activity in
windowsxp811203/Qwen3.8-27B-Abliterated
about 1 month ago
conv1d bloating issue on long contexts / dialogues
10
#2 opened about 2 months ago by
Qozimo
New activity in
MiniMaxAI/MiniMax-Music3
about 2 months ago
Frankenstein in Five Pieces: How to Build a Music Model from Spare Parts, Duct Tape, and Desperate Prayers
❤️
👍
3
38
#22 opened about 2 months ago by
AdrienneNoctis
New activity in
Qwen/Qwen3.8-2.4T-A95B
about 2 months ago
Huge disappointment: Qwen 3.8 open weights are text-only and stripped of Qwen 3.8 Max features (No Vision, No 1M Context)
👀
➕
26
26
#13 opened about 2 months ago by
NodeLinker
New activity in
AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
about 2 months ago
waiting for Aeon-qwen 3.8 27b -Ultimate-bf16
14
#10 opened 2 months ago by
Qozimo
New activity in
meta-models/Muse-Glimmer-30B
about 2 months ago
Even Big Lebowski doesn’t know where the money went: Meta’s "Superintelligence" is a sub-$5 basement-tier Frankenstein
🔥
18
17
#43 opened 2 months ago by
Qozimo
"The Future is for Everyone" — As Long as Everyone Has a Casio Calculator
🤯
👍
14
5
#48 opened 2 months ago by
Qozimo
New activity in
MiniMaxAI/MiniMax-H3
2 months ago
Where is the money, Lebowski? The "duct-tape" architecture review)))
➕
3
26
#46 opened 2 months ago by
Qozimo
New activity in
openpangu/openPangu-2.0-Pro
2 months ago
大厂PPT的神话,开源界的笑话:扒一扒 openPangu 2.0 里的“史诗级硬核套壳
👀
🤗
5
#2 opened 2 months ago by
Qozimo
New activity in
mistralai/Mistral-Medium-3.5-128B
5 months ago
Gemma 31B is your daddy: How to burn $1B and lose to a model 5x smaller+Huge thanks to the maintainers for closing the previous thread! wp nice one!
👍
1
1
#16 opened 5 months ago by
Qozimo
Gemma 31B is your daddy: How to burn $1B and lose to a model 5x smaller
👀
👍
6
8
#15 opened 5 months ago by
Qozimo
New activity in
mradermacher/model_requests
8 months ago
MMPROJ issue ; DavidAU/Gemma3-27B-it-vl-GLM-4.7-Uncensored-Heretic-Deep-Reasoning
❤️
👍
1
4
#1772 opened 8 months ago by
DavidAU
Load more