Optimized LLMs for edge devices using dynamic pruning, quantization, and distributed inference by making models fit where they shouldn't.
Research project optimizing Large Language Models for edge deployment. Implemented and evaluated techniques including dynamic pruning, quantization, and distributed inference to reduce computational requirements while preserving model quality on resource-constrained devices.
Key Features
Dynamic pruning for model size reduction
Quantization for efficient edge inference
Distributed inference across devices
Performance benchmarking and trade-off analysis
Technology Stack
AI/ML
PyTorchLLMQuantizationModel Compression
Tools
GitJupyter Notebooks
Challenges
Balancing model size, accuracy, and inference speed on edge hardware
Evaluating optimization techniques across different model architectures
Key Learnings
LLM compression and quantization techniques
Edge deployment constraints and optimization strategies
Research methodology for model performance evaluation