Federated learning (FL) enables privacy-preserving training on decentralized data but faces challenges on resource-constrained IoT devices due to high energy and communication costs. We address these by evaluating optimization strategies on a physical testbed of heterogeneous IoT devices. Our holistic approach, combining Top-K gradient compression with adaptive early stopping, reduces network bandwidth by 75% and lowers computational load without compromising accuracy. A compute-aware partitioning variant further rebalances workloads to minimize idle time. We identify the ”straggler” problem as a critical bottleneck, with faster devices idling for over 75% of the time. Finally, we provide a rigorous metrological assessment, quantifying Type A and Type B uncertainties to validate our findings. These findings highlight the need for system-aware optimizations for sustainable FL in IoT ecosystems.