OpenGVLab/Multi-Modality-Arena
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
Topics
Explore related topics
Jump into the topic listings this repository belongs to.
Join the conversation
Reviews · Questions · Posts
Share what you know about Multi-Modality-Arena — write a review from your real experience, ask an implementation question, or publish a post about how you use it.
Share your experience
Write or update your review
Explain what worked, what broke down, and what another team should know before adopting Multi-Modality-Arena.
Project Q&A
Questions and answers
Browse implementation threads tied directly to OpenGVLab/Multi-Modality-Arena. Each question links through to the full answer page.
Be the first to ask how teams run Multi-Modality-Arena in production. Every question you post becomes a durable, searchable answer page other developers can find.
Ask the first questionRelated posts
Posts tagged with the same topics
These posts come from the same topic surface as this repo, so readers can move from project evaluation into practical writeups and migration notes without leaving context.
Share how your team uses Multi-Modality-Arena — a migration note, an architecture writeup, or a comparison. Your post reaches everyone browsing these same topics.
Write the first post