Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Thilak Ruwan Chamara Denipitiya
ruwan
2
2
105
Follow
21world's profile picture
webxos's profile picture
Quazim0t0's profile picture
3 followers
·
51 following
ThilakDen
AI & ML interests
AI
Recent Activity
reacted
to
Enderchef
's
post
with 🤗
about 9 hours ago
Introducing GCI-Bench, Glint Research's first benchmark! This 5000-question bench benchmarks attention and gradients; it gives a prompt and unrelated filler parts. The goal is to benchmark if the model understands and attends to the important parts more than the unimportant parts. I've already benchmarked on the top 5 models on our Leaderboard(Glint2, Supra 50M, Glint1.3, GPT S 5M, GPT X2 125M) Pending benchmark status at https://huggingface.co/spaces/OpenEvals/README/discussions/2 https://huggingface.co/datasets/Glint-Research/GCI_Bench
reacted
to
FredyRivera-dev
's
post
with 🚀
about 18 hours ago
We wrote a full technical guide on how to train a bilingual (ES/EN) LLM from scratch: TinyQwen. Covers: - Hybrid architecture based on Qwen3.5 - Pre-training with 15B tokens - Cost benchmark between H200 and B200 - Post-training with SFT + LoRA - Full code and data, open source With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline. Blog post: https://aquiles-ai.vercel.app/blog/tinyqwen-from-scratch Technical feedback welcome, especially from anyone looking to replicate the pipeline with more compute.
liked
a model
1 day ago
nightmedia/Qwen3.5-9B-Holodeck-Fara
View all activity
Organizations
None yet
models
5
Sort: Recently updated
ruwan/MyGemmaNPC
Text Generation
•
0.3B
•
Updated
Aug 24, 2025
•
5
ruwan/story_model_final
16.2M
•
Updated
Mar 25, 2025
•
1
ruwan/llama-tiny
Text Generation
•
Updated
Dec 11, 2023
•
6
ruwan/open-llama-sharded-1GB-7B-alpaca-vmware
Text Generation
•
Updated
Jun 8, 2023
•
6
ruwan/open-llama-sharded-3GB-7B-alpaca-vmware
Text Generation
•
Updated
Jun 8, 2023
•
4
datasets
0
None public yet