0% Complete
فارسی
Home
/
شانزدهمین کنفرانس بین المللی فناوری اطلاعات و دانش
An LLM-Based Approach for Clarifying the Decisions of Vision Models in Autonomous Vehicles
Authors :
Omid Mosalmani
1
Mohammad Javad Rashti
2
Seyed Enayat Alavi
3
1- دانشگاه شهید چمران اهواز
2- دانشگاه شهید چمران اهواز
3- دانشگاه شهید چمران اهواز
Keywords :
Explainable AI،Prompt Engineering،Large Language Models،Autonomous Vehicles،Textual Explanation
Abstract :
With the increasing utilization of autonomous vehicles, the transparency and explainability of their decisions have become crucial for gaining user trust and enhancing road safety. Current textual explanation methods rely on limited datasets, leading to repetitive and superficial explanations. This research presents a hybrid system where the ADAPT decision-making model is used to predict driving actions, and its attention maps serve as an interface between visual data and the explanation module. Subsequently, large language models, from the Gemini and GPT families, receive the final decision, the attention map, and a carefully designed prompt to generate concise and understandable textual explanations. The primary innovation of this approach lies in combining the decision-making model with LLMs, leveraging their extensive knowledge beyond the constraints of training data to enable the generation of more precise and diverse explanations. The system is evaluated on the BDD-X dataset and measured against standard captioning metrics including BLEU-4, METEOR, ROUGE-L, CIDEr-D, and SPICE. The evaluation results indicate the superiority of explanation outputs in our system, compared to the baseline ADAPT, particularly in multi-reference scenarios, providing more fluent and contextually rich explanations. For instance, the output acquired from Gemini 2.5 Pro model achieves a METEOR score of approximately 19.45, a significant improvement of about 28 percent compared to 15.2 for ADAPT. Furthermore, supplementary experiments show that using a contour representation of the attention map and fine-tuning the models lead to increased visual-textual consistency and result stability. In summary, by linking the visual attention of the decision-making model to the linguistic capabilities of LLMs, this research takes a step toward developing more explainable and trustworthy autonomous vehicles.
Papers List
List of archived papers
Knowledge Graph Based Retrieval-Augmented Generation for Multi-Hop Question Answering Enhancement
Mahdi Amiri Shavaki - Pouria Omrani - Ramin Toosi - Mohammad Ali Akhaee
LuckyAgent2022: A Stop-Learning Multi-Armed Bandit Automated Negotiating Agent
Arash Ebrahimnezhad - Faria Nassiri-Mofakham
Design and Simulation of an Accident Prevention System Based on Weather Conditions and Internet of Things
Forouzan Dastbaz - Abdolah Chalechale
سیستم پیشنهاددهنده غذای سالم با استفاده از داده کاوی عادت های تغذیه ای کاربران
محمد عباسی - مریم حسینی پزوه - محمدرضا شمس
Wireless Virtual-Reality by considering Hybrid Beamforming in IEEE802.11ay standard
Nasim Alikhani - Abbas Mohammadi
Classification of Personality Traits on Facebook Using Key Phrase Extraction, Language Models and Machine Learning
Faezeh Safari - Abdolah Chalechale
آسیب شناسی استقرار بلاکچین در صنعت بانکی کشور ایران
نیلوفر مرادحاصل
قطعه بندی خودکار توده کلیه در تصاویر توموگرافی کامپیوتری با استفاده از همافزایی شبکه عصبی عمیق U-Net و الگوریتم فراابتکاری نهنگ
علی خلیلی - محمد مصلح - محمد خیراندیش
رویکرد نوین مبتنی بر خوشهبندی محلی شدت روشنایی برای جداسازی بافتهای مغزی
آسیه خسروانیان - سعید آیت
نظرکاوی در سطح مفهوم با استفاده از رویکردی ترکیبی
سیدرضا قادریان خیرآبادی سیدرضا قادریان خیرآبادی -
more
Samin Hamayesh - Version 43.8.0