50 results
Open links in new tab
  1. VisLang - Vision, Language and Learning Lab at Rice University

    Research group at Rice University led by Vicente Ordonez, working at the intersection of computer vision, natural language …

  2. vislang.ai

    @inproceedings{smith2023construct, title = {ConStruct-VL: Data-Free Continual Structured VL Concepts Learning.}, author = {Smith, …

  3. MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible …

    Universal multimodal embedding models have achieved great success in capturing semantic relevance between queries and …

  4. Learning to Name Objects - VisLang Lab

    Learning to Name Objects by Vicente Ordonez, Wei Liu, Jia Deng, Yejin Choi, Alexander C. Berg, Tamara L. Berg.. Communications …

  5. Chat-crowd: A Dialog-based Platform for Visual Layout Composition

    본 논문에서 우리는 대화형 상호작용을 통한 시각적 레이아웃 구성을 위한 인터랙티브 환경인 Chat-crowd를 소개한다. Chat-crowd는 두 …

  6. On the Transferability of Visual Features in Generalized Zero-Shot ...

    Generalized Zero-Shot Learning (GZSL) nhằm huấn luyện một bộ phân loại có thể khái quát hóa sang các lớp chưa thấy, sử dụng …

  7. Gender Bias in Coreference Resolution: Evaluation and Debiasing …

    Wir führen einen neuen Benchmark, WinoBias, für die Koreferenzauflösung mit Fokus auf geschlechtsbezogene Verzerrungen ein. …

  8. One Model, Many Budgets: Elastic Latent Interfaces for Diffusion ...

    Diffusion transformers (DiTs) achieve high generative quality but lock FLOPs to image resolution, limiting principled latency-quality …

  9. FlexEControl: Flexible and Efficient Multimodal Control for Text-to ...

    Các mô hình khuếch tán text-to-image (T2I) có thể điều khiển tạo ra các ảnh được điều kiện trên cả lời nhắc văn bản lẫn các đầu …

  10. Learning from Synthetic Data for Visual Grounding - VisLang Lab

    이 논문은 텍스트 설명을 이미지 영역에 그라운딩하는 비전-언어 모델의 능력을 향상시키기 위한 합성 학습 데이터의 효과를 광범위하게 …

  11. Publications - VisLang Lab at Rice University

    Academic publications from the VisLang Lab at Rice University on computer vision, natural language processing, and machine learning.