Showing posts with label multimodal. Show all posts
Showing posts with label multimodal. Show all posts

Thursday, November 02, 2023

On Ferret: Refer and Ground Anything Anywhere at Any Granularity

Just finished reading this paper. Very impressive work!

"We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with 95K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination. Code and data will be available at this https URL"

[2310.07704] Ferret: Refer and Ground Anything Anywhere at Any Granularity

Wednesday, January 06, 2021

OpenAI debuts DALL-E for generating images from text

Recommendable! Text to image synthesis! Multimodal machine learning is imminent!

"... Tests OpenAI shared today appear to demonstrate that DALL-E has the ability to manipulate and rearrange objects in generated imagery and also create things that don’t exist, like a cube with the texture of a porcupine or a cube of clouds. Based on text prompts, images generated by DALL-E can appear as if they were taken from the real world or can depict works of art. ..."

"... We’ve found that it has a diverse set of capabilities, including creating anthropomorphized versions of animals and objects, combining unrelated concepts in plausible ways, rendering text, and applying transformations to existing images. ..."

OpenAI debuts DALL-E for generating images from text | VentureBeat

Here is the link to the respective OpenAI blog post:
DALL·E: Creating Images from Text We’ve trained a neural network called DALL·E that creates images from text captions for a wide range of concepts expressible in natural language



Thursday, December 31, 2020

The immense potential and challenges of multimodal AI

Recommendable! One of the latest trends in AI: Combining text, video, audio, and images etc.

"... While systems capable of making these multimodal inferences remain beyond reach, there’s been progress. New research over the past year has advanced the state-of-the-art in multimodal learning, particularly in the subfield of visual question answering (VQA), a computer vision task where a system is given a text-based question about an image and must infer the answer. As it turns out, multimodal learning can carry complementary information or trends, which often only become evident when they’re all included in the learning process. ...
In multimodal systems, computer vision and natural language processing models are trained together on datasets to learn a combined embedding space, or a space occupied by variables representing specific features of the images, text, and other media. ..."

The immense potential and challenges of multimodal AI | VentureBeat