[SenSys '26] PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling
-
Updated
Aug 21, 2026 - Python
[SenSys '26] PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling
Focuses on running LLMs at the edge (on devices like Raspberry Pi). Why it works: Highlights the project’s edge-computing nature and AI capabilities.
A high-performance, modular AI chat solution for Jetson™ edge devices. It integrates Ollama with the Meta Llama 3.2 3B model for LLM inference, FastAPI-based Langchain middleware, and OpenWebUI.
To associate your repository with the edgellm topic, visit your repo's landing page and select "manage topics."