Abstract
Effective model deployment on low-resource devices is becoming more crucial as edge artificial intelligence applications get more complicated. Pruning, quantization, and knowledge distillation are potential automated model compression methods because they minimize model size and computation without compromising accuracy. whether alternative model compression strategies improve artificial intelligence models for edge devices like IoT sensors, mobile devices, and embedded systems. These solutions enable real-time inference and reduce latency by dynamically changing models to hardware restrictions. Automatic compression does this. We study how different compression schemes affect model performance, power consumption, and memory usage to determine whether they are feasible for edge AI applications. The challenges of balancing compression and accuracy and ensuring interpretability in compressed models are also examined. the development of scalable and efficient edge artificial intelligence systems to accelerate AI deployment in various settings with limited resources.

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
