智慧育樂獎 | 財團法人工業技術研究院 | GAI生成式多模態智慧顯示系統

得獎方案簡介    
本項目旨在重新定義智慧育樂與導覽模式。系統以大型語言模型(LLM)為跨模態語意中樞,深度整合語音辨識、事件式影像辨識與追蹤、生成式影像技術,建構生成式多模態智慧顯示代理系統。透過語意理解驅動即時影像追蹤與內容生成,使導覽由被動呈現化為即時主動回應。 已於壽山動物園黑熊與白犀牛展區實場營運,成功解決戶外導覽動物不可視及導覽人力不足的痛點,實現自然語言互動、高準確率的動態追蹤與生成式互動影像,將傳統顯示設備升級為具備決策能力的 AI 互動代理,重新定義智慧育樂模式。

 

得獎方案創新特色

以 LLM 為語意決策核心,整合語音辨識、多攝影機追蹤與生成式影像技術,實現結合「即時理解、動態追蹤、生成回應」之智慧顯示應用,提升展示互動性與教育深度。

  • 智慧語意中樞:利用LLM達成90%以上語意準確率,精準解析口語化指令並即時協調多模態AI模組運作。
  • 高精度空間追蹤:整合2D影像與3D地形資訊,將定位精度從1.5m大幅提升至0.5m,有效解決開放場域之遮蔽與視角受限問題,獲得連續即時追蹤畫面。
  • 高一致性生成式角色:結合Stable Diffusion與LoRA微調技術,確保生成之虛擬角色在不同情境下的視覺一致性達99%。

 

得獎方案應用

本方案不僅為單一導覽系統創新,更代表顯示產業價值鏈的關鍵升級。透過導入 LLM 與生成式 AI,將傳統顯示設備由硬體供應轉型為具備語意理解與互動能力的智慧系統平台,開啟「顯示 × AI × 服務」的新產業模式。 系統已於壽山動物園完成實場驗證,並成功技術移轉予達擎股份有限公司,具備跨場域複製能力,可延伸應用於國內外文教育樂及觀光等場景。

 


Smart Display Application Awards - Smart Edutainment Award

Company Name: Industrial Technology Research Institute (ITRI)

Technology/Product Name:GAI-Based Multimodal Intelligent Display System

 

Solution Features

This project redefines smart edutainment using LLMs as a semantic hub to integrate speech, event-based vision, and generative AI. The system creates a "Smart Display Agent" that transforms passive tours into proactive interactions. Deployed at Shoushan Zoo for black bears and rhinos, it solves animal invisibility and staffing shortages through real-time tracking and natural language interaction, upgrading traditional displays into decision-making AI agents.

 

Innovation

Using LLMs as the semantic decision core, this project integrates speech recognition, multi-camera tracking, and generative AI to enable smart displays with real-time understanding, dynamic tracking, and generative responses, enhancing interactive engagement and educational depth.

  • Intelligent Semantic Core: Optimized intent recognition to 90%+ accuracy using LLM to parse natural language and coordinate multimodal AI modules in real-time.
  • High-Precision Spatial Tracking: Optimizes positioning to 0.5m by integrating 2D/3D data, overcoming environmental occlusions.
  • Consistent Visual Generation: Pioneered the use of Stable Diffusion with LoRA fine-tuning to ensure 99% visual consistency for generated personas across diverse scenarios.

 

Applications

This solution represents a critical upgrade of the display industry value chain. By integrating LLMs and Generative AI, it transforms traditional hardware into an intelligent platform with semantic understanding, pioneering a new "Display × AI × Service" industrial model. Following successful field validation at Shoushan Zoo, the technology has been officially transferred to AUO Display Plus (ADP). With its high scalability, this system can be replicated globally across edutainment, cultural, and tourism sectors.

(Image)
  • 得獎年分

    2026
  • 作品名稱

    GAI生成式多模態智慧顯示系統
  • 得獎公司

    財團法人工業技術研究院
  • 公司官網

    前往連結
TOP