SkyAdvisor: A Hybrid Large Language Model and Contextual Bandit Framework for Conversational Flight Recommendations
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE credits
Student thesis
Abstract [en]
Introduction: Recommender systems are widely used in modern flight comparison platforms, yet many rely on offline ranking models that cannot adapt to individual user preferences during an ongoing search session. This thesis investigates SkyAdvisor, a hybrid conversational flight recommendation framework that combines an offline learning-to-rank model with an online contextual bandit to support adaptive flight recommendations based on implicit user feedback.
Research Question: The research question addressed in this thesis is to what extent integrating an offline learning-to-rank model with a contextual bandit affects the performance of flight recommendations compared to a contextual bandit operating alone.
Method: To answer this question, two system architectures were designed and, due to the absence of a live deployment environment, evaluated through controlled offline simulation across 30 repeated trials of 10,000 interaction rounds each. System A used a contextual bandit alone as the only ranking component. In contrast, System B first applied a learning-to-rank model to pre-rank candidate flights before passing them to the contextual bandit to select the best option. Both systems used a large language model as a conversational interface to extract structured user preferences from natural-language input, including travel preferences. The systems were evaluated using cumulative reward, final click-through rate, cumulative regret, and Precision@1.
Results: The results show that the hybrid architecture improved reward-based performance for the contextual bandit when it learned entirely from online interactions, with significantly higher cumulative reward and click-through rate than the bandit-only baseline. However, the improvement was not consistent across all metrics, as cumulative regret and the proportion of optimal selections did not improve significantly. The results also show that when the bandit was initialized using historical interaction data, its performance could degrade if that historical data did not match the distribution of candidates produced by the pre-ranking stage.
Discussion: The thesis concludes that combining learning-to-rank with contextual bandits can improve flight recommendation performance, but the benefit is conditional rather than universal. A hybrid architecture is most effective when the offline ranker improves the quality of the candidate pool without creating a distribution mismatch for the online bandit component.
Place, publisher, year, edition, pages
2026.
Keywords [en]
Contextual Bandits, Learning-to-Rank (LTR), Conversational Recommender Systems, Flight Recommendation, Implicit Feedback, Warm- Start Initialization, Exploration–Exploitation
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:su:diva-257323OAI: oai:DiVA.org:su-257323DiVA, id: diva2:2079930
2026-06-252026-06-25