ED Mathématiques et Informatique
Scheduling algorithms for the optimization of distributed machine learning models on heterogeneous resources
by Alan LIRA NUNES (LaBRI - Laboratoire Bordelais de Recherche en Informatique)
The defense will take place at 15h00 - 310 Universidade Federal Fluminense (UFF), Instituto de Computação, Rua Passo da Pátria, 156, - São Domingos, Niterói, RJ, 24210-310, Brésil
in front of the jury composed of
- Laércio LIMA PILLA - Chargé de recherche - Université de Bordeaux - Directeur de these
- Shadi IBRAHIM - Chargé de recherche - INRIA - Rapporteur
- Alfredo GOLDMAN - Professeur des universités - Universidade de São Paulo - Rapporteur
- Valmir BARBOSA - Professeur des universités - Universidade do Estado do Rio de Janeiro - Examinateur
- César DE ROSE - Professeur des universités - Pontifícia Universidade Católica do Rio Grande do Sul - Examinateur
- Débora SAADE - Professeure des universités - Universidade Federal Fluminense - Examinateur
- Lúcia DRUMMOND - Professeure des universités - Universidade Federal Fluminense - CoDirecteur de these
Federated Learning (FL) enables multiple clients to collaboratively train machine learning models without sharing their raw data. However, practical cross-device FL remains challenging because clients differ in computational capacity, communication quality, available energy, local data distributions, and reliability. In synchronous FL, these differences affect system efficiency and learning performance: slow, energy-constrained, or poorly representative clients can delay training, increase resource consumption, or degrade convergence. Most existing strategies only decide whether a client participates in a communication round, while selected clients usually train the model on all their local data, limiting per-client workload adaptation. This thesis addresses client selection in synchronous cross-device FL from a task scheduling perspective. Clients are modeled as heterogeneous resources, and the workload of a round is decomposed into tasks corresponding to subsets of local data. Thus, the server can decide which clients participate and how much data each one should process. This formulation enables fine-grained workload allocation over heterogeneous resources and supports the joint optimization of system-level and learning-related objectives. This thesis first proposes MEC and ECMTC, two optimal scheduling algorithms for workload allocation considering time and energy. MEC first minimizes round duration and then energy consumption, whereas ECMTC first minimizes energy consumption and then round duration under a time constraint. These algorithms rely on dynamic programming and provide optimal schedules for their respective objective orderings. The thesis then introduces MetaCS-FL, a metaheuristic-based framework that extends this perspective to a broader multi-objective setting. Beyond execution time and energy, MetaCS-FL considers model utility, class-distribution quality, participation diversity, and client reliability. It uses the schedule produced by ECMTC as an initial solution and refines task assignments to clients through the Large Neighborhood Search metaheuristic. The framework also integrates event-driven reselection, solution reuse, reliability-aware capacity control, and differential privacy, allowing noisy information about class distributions to be exploited without accessing exact distributions. The proposed methods are evaluated in image and text classification, under IID and non-IID distributions, using heterogeneous emulated clients and comparisons with FedAvg and state-of-the-art strategies. In the evaluated scenarios, MetaCS-FL outperforms the other algorithms overall and achieves the best trade-off among training time, energy, convergence speed, and fairness. Under static availability, it reduces total training time and energy consumption while reaching the target accuracy in fewer rounds, without concentrating the workload on few clients. The private variant remains close to the non-private version, indicating that noisy disclosure of distributions preserves useful selection information while improving privacy. Under dynamic availability, with late-joining or intermittent clients, MetaCS-FL maintains superior performance by adapting workload assignments according to client availability, completion behavior, and reliability.