Releasing smolbenchmark: Helps you choose the best model for your hardware!
Reddit r/LocalLLaMA3d4 min read
Most model leaderboards assume a server with powerful GPUs to run models that people daily use. However, my smolbenchmark is the other column: models that fit in 8GB, ranked by: decode speed, tokens per joule, and heat, and all of this on your OWN hardware ranging from: tablets phones macs jetsons raspberry pis Currently, 13 families on the chart right now, ~1000 configs for the Jetson nano Orin Super 8GB. One device is live measuring: tok/s tok/J ITL latency power metrics thermals and battery Models that are small enough to actually fit on a device that you own. All the performance benchmarki