WebsiteHunt is broughtBrought to you by Redact Everything
FM
Flash-MoE: 397B on a Laptop
Running a big model on a small laptop.
Flash-MoE runs a 397B Mixture-of-Experts model on a laptop with a pure C/Metal inference engine. Streams a 209GB model from SSD with 4-bit quantization and no Python for production-quality output.
Discussion highlights
AI-generated summary based on Hacker News comments
🔗 Full discussion: https://news.ycombinator.com/item?id=47476422
Comments
Loading comments...
Category
You might also like
Labrynth AI
AI regulatory intelligence platform 🌱 Sponsored
ProjectManagementTools
Optimize your Projects with ProjectManagementTools 🌱 Sponsored
Introducing Claude Opus 5
Anthropic's model for long-running agents and coding.
CA
Carton
Run any ML model from any programming language.
Smarter Quiz
Are you smarter than a language model?
Project Vend
Exploring AI's role in managing a small shop
MF
Mavericks Forever
A guide to running Mac OS X Mavericks
Tokenware
One API for All AI Models
Needle: Tiny on-device AI
Tiny, open-source AI that runs on small devices.
OS
Operating System in 1,000 Lines
Build a small operating system from scratch
GC
Grok Code Fast 1
A speedy and economical reasoning model for coding.
Voicebox
An all-in-one generative Al model for speech
DeepSeek V4 Pro 0813
GA release of a large-scale mixture-of-experts model
Replit Code V1.5 3B
A new code generation language model
Nativ
Run AI locally on Apple Silicon.
JigsawStack
Small models that power your tech stack
YE
You Exist In The Long Context
Thoughts on long-context AI models by Steven Johnson
Gradio
demo your machine learning model
LLaMA
A foundational, 65-billion-parameter large language model
Talkie: Vintage 13B Language Model
Exploring AI history by training models on pre-1931 text.
PrismML Bonsai 27B: On-device
First 27B-class model that runs on a phone with multimodal reason