Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Luca Modica
lucamodica
6
18
Follow
Benson's profile picture
EvilScript's profile picture
2 followers
·
11 following
AI & ML interests
None yet
Recent Activity
reacted
to
SeaWolf-AI
's
post
with 🔥
about 11 hours ago
A small gift for anyone building or studying foundation models. Most "open" models hand you the weights and stop there. With Aether-7B-5Attn we wanted to hand over the whole thing — so you can actually learn from it, reproduce it, and build on it: the data recipe, the training code, every hyperparameter, the complete logs, and the intermediate checkpoints. All Apache-2.0, reproducible byte-for-byte. What you can do with it: 🔁 Rebuild it from scratch, or fork the recipe for your own model 🔬 Study a real heterogeneous-attention MoE — 49 layers place 5 attention mechanisms on a 7×7 Latin square, arranged as a clean, attributable ablation 📈 Trace training dynamics across the released checkpoints (110k / 115k / 162k) It's a modest 6.59B model, and an honest one — the limitations (no KV-cache in this build, small scale) are written right in the card. We're not claiming it's special. If any piece of it saves you time or teaches you something, that's exactly what we hoped for. 🤗 📖 Full write-up → [blog] · https://huggingface.co/blog/FINAL-Bench/opensource-llm 📦 5 Attention Base · https://huggingface.co/FINAL-Bench/Aether-7B-5Attn 🎯 5 Attention Instruct · https://huggingface.co/FINAL-Bench/Aether-7B-5Attn-it 🚀 5 Attention Live demo · https://huggingface.co/spaces/FINAL-Bench/Aether-Sovereign-AI 📦 7 Attention Base · https://huggingface.co/FINAL-Bench/Aether-7B-7Attn-base 📦 11 Attention Base · https://huggingface.co/FINAL-Bench/Aether-6B-11Attn-base 🧬 Collection · https://huggingface.co/collections/FINAL-Bench/aether-foundation-model #opensource #LLM #MoE #reproducibility #Apache2
upvoted
a
collection
3 days ago
Alpamayo
upvoted
a
paper
10 days ago
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
View all activity
Organizations
lucamodica
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
2 datasets
3 months ago
yaak-ai/L2D
Viewer
•
Updated
May 26
•
26.5M
•
60k
•
49
Qwen/RationaleRM
Preview
•
Updated
Feb 5
•
435
•
28
liked
a model
6 months ago
Qwen/Qwen3-VL-Embedding-2B
Sentence Similarity
•
2B
•
Updated
Apr 16
•
1.25M
•
•
434
liked
a dataset
6 months ago
ivc-lrp/STSBench
Viewer
•
Updated
May 16, 2025
•
971
•
37
•
5
liked
2 datasets
7 months ago
aaaaaap/unstructed
Updated
Oct 26, 2025
•
762
•
14
turing-motors/CoVLA-Dataset
Viewer
•
Updated
Mar 1, 2025
•
10k
•
4.09k
•
52
liked
a model
8 months ago
nvidia/Alpamayo-R1-10B
Robotics
•
11B
•
Updated
13 days ago
•
18.9k
•
420
liked
a dataset
8 months ago
Telkwevr/Bench2Drive-VL-base
Updated
Apr 7
•
1.19k
•
1
liked
a dataset
9 months ago
rethinklab/Bench2Drive-Full
Updated
Jul 22, 2024
•
5.87k
•
7
liked
3 models
9 months ago
exiawsh/OmniDrive
Updated
Apr 18, 2025
•
4
exiawsh/pretrain_qformer
Updated
Jun 9, 2025
•
621
•
2
allenai/MolmoAct-7B-D-0812
Robotics
•
8B
•
Updated
Oct 24, 2025
•
924
•
53
liked
a model
10 months ago
ryota-komatsu/s5-hubert
92M
•
Updated
Nov 6, 2025
•
56
•
1
liked
a dataset
11 months ago
McAuley-Lab/Amazon-Reviews-2023
Updated
Dec 8, 2024
•
77.7k
•
325
liked
a model
12 months ago
myshell-ai/MeloTTS-English
Text-to-Speech
•
Updated
Dec 24, 2024
•
160k
•
314
liked
2 models
about 1 year ago
zai-org/glm-4-voice-9b
10B
•
Updated
Oct 25, 2024
•
17.3k
•
119
PlayHT/PlayDiffusion
Updated
Jul 29, 2025
•
111
liked
a dataset
almost 3 years ago
lambda/pokemon-blip-captions
Viewer
•
Updated
Mar 20, 2024
•
833
•
1.43k
•
308