The workshop is a half-day event on August 17th, running in two 90-minute blocks (14:00 – 15:30 and 16:00 – 17:30) separated by the conference coffee break, and combining an invited keynote, contributed paper presentations, and an interactive discussion round. The workshop will take place in Room SFG-2030.

TimeSession
14:00 – 14:05Welcome & Introduction
14:05 – 14:40Keynote
Prof. Dr. Norbert Wehn 🌐 β€” RPTU University, Germany
"It’s all about Energy Efficiency - A memory perspective on AI"
30 min talk + 5 min Q&A
14:40 – 15:30Paper Session I β€” Power Usage
Three full-paper presentations (14 min talk + 3 min Q&A each)
  1. Safety-Constrained Contextual Bandit for Dynamic Power Management
  2. WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs
  3. WattLayer: Get Layers Right to Estimate Inference Energy of Neural Networks
15:30 – 16:00β˜• Coffee Break
16:00 – 16:50Paper Session II β€” Efficient Models & Deployment
Three full-paper presentations (14 min talk + 3 min Q&A each)
  1. CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
  2. TinyHybrid: Ultra-Efficient CNN-Transformer Architecture for Edge-Deployable Brain Tumor Classification
  3. ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
16:50 – 17:15Short Paper Session
Two short-paper presentations (10 min talk + 3 min Q&A each)
  1. Efficient Vision Models for Jetson: Steel Classification via Knowledge Distillation
  2. Accounting for Bias Enables Sustainable LLM Evaluation
17:15 – 17:25Position Paper & Discussion
Short pitch (3 min) followed by an open audience discussion (7 min)
17:25 – 17:30Closing Remarks
including the Best Paper Award πŸ†