UDOP
UDOP adopts an encoder-decoder Transformer architecture based on T5 for document AI tasks like document image classification, document parsing and document visual question answering. You can use the model for document image classification, document parsing…
Structured knowledge from LLMs improves prompt learning for visual language models
Using LLMs to create structured graphs of image descriptors can enhance the images generated by visual language models. Learn how structured knowledge can improve prompt tuning for both visual and language comprehension.
AI, Cognition and the Economy
AICE establishes a global research network dedicated to cultivating an interdisciplinary research community. This collective will delve into the profound impact of Generative AI (GAI) on human cognition, work dynamics, and economic growth. This pioneering…