OpenArt, the four-year-old San Francisco creative AI tooling startup co-founded by former Googlers Coco Mao and John Qiaois, wants to answer an increasingly pertinent question for enterprises and individuals using AI models to create multimedia: not simply “which model is best?” but “which model is best for the specific media job I need to complete?”

That's why today, the firm is launching a new public benchmark for AI image and video generation: the brand new OpenArt Arena divides model rankings into use-case boards for work including filmmaking, e-commerce, graphic design, motion design, video editing and lip sync, using blind pairwise evaluations from creative practitioners and a larger pool of what the company calls “tastemakers.”

The company says it will publish its methodology and part of its prompt set at launch, while keeping some active evaluation prompts private to reduce the risk of models being tuned to the benchmark.

The idea is straightforward but increasingly necessary. Image and video generation models now ship quickly enough that a creative team can spend substantial time simply deciding whether to use OpenAI’s GPT Image series, Google’s Nano Banana family, ByteDance’s Seedream and Seedance models, Alibaba’s Wan, xAI’s Grok Imagine or another option.

“Our goal is for this to become a new standard, for more and more people to know it and recognize it," Stella Guan, Head of Growth & Operations at OpenArt AI, told VentureBeat in an exclusive interview conducted before the Arena’s final voting was completed. "When creative people want to pick the best image or video model they should use, they would think about OpenArt Arena and know that it is the most credible industry standard for them to pick the model.”

OpenArt itself is built around that abundance: its platform lets users work across major third-party image and video models and also sells its own higher-level creation products, including Director, its conversational “vibe directing” workflow for generating multi-scene videos.

OpenArt’s current model catalog includes image, video, audio and 3D systems from a range of providers.

“Which is the model should I pick?” is one of the questions OpenArt gets repeatedly from enterprise customers, Guan said.

For companies coming from traditional industries and only now adopting generative AI, she said, picking a model can be “the first challenge that they have when it comes to adopting AI.”

But how do AI video creators and artists pick the best model at present?

One prolific and pseudonymous creator, @BLVCKL!GHT, told VentureBeat via X direct message: "Which model I use depends entirely on the work required. Looping animations and more dynamic scenes call for different tools, so a project might land on Minimax or Seedance 2.5 based on that alone. Clients often ask for state of the art models, which makes the choice for me, but budget ultimately shapes what's practical for any given project."

Organized around specific creative AI media jobs

Instead of publishing one image score and one video score, OpenArt is building separate boards around jobs and creative disciplines.

This includes separate leaderboards showing how AI image and videos models rate (according to OpenArt's handpicked creative professional users and outside experts, numbering more than 800) on completing tasks in animation, advertising, graphic design, e-commerce, film, motion design, video editing, lip sync and overall.

Each board is supposed to use criteria tailored to the work. Guan gave filmmaking as one example: camera movement, lighting, cinematic quality and realistic skin can matter more there, while an advertising workflow may put more weight on readable text, accurate logos, product fidelity and placement.

That design matters because a model that is strong at one specific creative task is not necessarily the best at another.

The distinction is more consequential for businesses than it may initially sound. A marketing department generating packaging shots, an internal studio previsualizing a commercial, and a post-production team altering existing footage may all describe themselves as “using AI video” while placing very different demands on the underlying model.

A leaderboard that collapses those jobs into one preference score can be useful for finding broadly capable systems while still giving an enterprise team little guidance on which model to route a particular production task to.

OpenArt says it generates outputs from curated prompt sets for each board and presents them to evaluators without model labels.

Judges make side-by-side choices rather than assigning isolated scores, and OpenArt aggregates those preferences using the 1952 Bradley–Terry statistical model, an equation that calculates an item's hidden probability of winning based on its head-to-head performance against competitors.

The judging pool has two layers. The smaller Creative Expert Council includes named practitioners such as Emmy-winning animation director William Lau, creative technologist Willonius “King Willonius” Hatcher, marketing leader David Shing, and executives and educators from organizations including Edelman and UCLA.

Guan said OpenArt was also recruiting roughly 800 to 1,000 “tastemakers” from OpenArt users, outside creative communities and working professionals in fields such as advertising.

That number should be read as the planned pool, not a verified count of people who completed launch evaluations. OpenArt did not provide the final number of judges who voted, the total number of pairwise judgments, or the number of prompts in each launch benchmark in the materials reviewed for this story.

Guan said the company intends to publish a portion of each prompt set and may release older prompts after future refreshes, but not the entire active set. “One reason that we don’t publish everything at launch is that it would be easy to fit toward those data sets,” she said.

That is a legitimate benchmark-design concern: a fully public static test can eventually measure how effectively model developers optimize against the test itself rather than how well their systems generalize to unseen creative work.

But withholding test material creates a corresponding transparency burden. Users have to trust that the undisclosed prompts adequately represent the profession and that model outputs are generated, sampled and compared consistently.

What the first OpenArt Arena rankings show

The video results provide the clearest headline from OpenArt Arena’s launch rankings.

ByteDance’s Seedance 2.5, a proprietary, API-served model for which ByteDance has not released public weights, leads the overall video board with a score of 1,081, followed by Alibaba’s Wan 3.0 at 1,004 and ByteDance’s likewise closed-weight Seedance 2.0 at 1,000.

OpenArt Arena video overall leaderboard rankings

OpenArt Arena video overall leaderboard rankings. Credit: OpenArt

Wan 3.0 is the outlier on openness: Alibaba describes it as open source and has published an Apache 2.0-licensed repository, although that repository currently contains documentation and licensing rather than downloadable model weights, so it should not yet be treated as a conventionally open-weight, self-hostable release.

Seedance 2.5’s advantage carries across most of the more specialized video categories OpenArt evaluated: of the five video leaderboards supplied for the launch, it ranks first on four and second on the fifth. On the film leaderboard, Seedance 2.5 takes first at 1,049, ahead of Seedance 2.0 at 1,000 and Wan 3.0 at 958. It also leads motion design at 1,059, followed by Seedance 2.0 at 1,000 and Wan 3.0 at 990.

OpenArt arena video editing leaderboard rankings

OpenArt arena video editing leaderboard rankings. Credit: OpenArt

Video editing is the exception (barely) — and one of the more useful results for illustrating OpenArt’s argument that model selection should depend on the job.

Wan 3.0 ranks first at 1,034, Seedance 2.5 follows at 1,033, with Google’s proprietary Gemini Omni Flash third at 1,021.

Google distributes Omni Flash as part of its Gemini model family rather than as downloadable open weights. Seedance 2.5 then returns to the top for lip sync, scoring 1,063, ahead of Seedance 2.0 at 1,000 and Seedance 2.0 Mini, another hosted ByteDance variant rather than an open-weight release, at 994.

In other words, Seedance 2.5 appears to be the strongest general video performer in OpenArt’s launch testing, but it does not nominally win every specialized task.

OpenArt displays 95% confidence intervals on the boards, an important qualification when scores are close. A one-point difference such as Wan 3.0’s 1,034 versus Seedance 2.5’s 1,033 in video editing should not be interpreted as evidence of a meaningful quality gap by itself.

The dominance of Seedance 2.5 is reflected in its usage across the AI creator space. Zack London, better known as the high-quality, sci-fi/fantasy AI video creator Gossip Goblin, who has a feature length AI film premiering in hundreds of theaters later this year entitled Gods Don’t Give Gifts, told VentureBeat via direct message on X:

"A year ago, the market was pretty diverse — Hailuo, Kling and Veo all had their merits, along with more niche lip-sync tools like Sync Labs, Hedra and HeyGen. Now Seedance is so profoundly far ahead of everything else that it’s not even a discussion. I can say with confidence that none of the reputable creators in this space use anything but Seedance anymore. Hailuo’s H3 is being leveraged for mass spam and automation, but for serious work, the only option is Seedance."

For enterprise buyers, the more useful signal is the pattern across boards — which models repeatedly perform strongly, where changing the production task changes the ordering, and whether a model can be self-hosted or instead requires dependence on a vendor or API provider.

The image rankings are more fragmented, with no single model sweeping the three launch boards supplied for this story.

OpenArt Arena graphic design leaderboard rankings

OpenArt arena video editing leaderboard rankings. Credit: OpenArt

OpenArt Arena image editing leaderboard rankings

OpenArt Arena image editing leaderboard rankings. Credit: OpenArt

OpenAI’s proprietary GPT Image 2, available through OpenAI’s API rather than as an open-weight model, ranks first for graphic design and image editing, in both cases, ranking at 1,000.

Yet, when it comes to film-oriented imagery, Seedream 5.0 Pro leads at 1,014, while Google’s closed-weight Nano Banana Pro (Gemini 3 Pro Image) and GPT Image 2 both post scores of 1,000 in second and third place, respectively. Seedream 5.0 Pro also tops the e-commerce image board at 1,004, narrowly ahead of GPT Image 2 at 1,000 and Google’s similarly proprietary Nano Banana 2 (Gemini 3.1 Flash Image) at 993.

Overall, when it comes to image models performance across all tasks, OpenArt's evaluators preferred Alibaba's Seedream 5.0 Pro with a score of 1,010, besting GPT Image 2 by 10 points. Note that OpenAI updated the model to GPT-Images-2.5 last week, meaning OpenArt's leaderboard will also need to be updated to account for the change.

OpenArt Arena overall image model leaderboard rankings

OpenArt Arena overall image model leaderboard rankings. Credit: OpenArt

The openness distinction is therefore fairly stark among OpenArt’s launch leaders:Wan 3.0 is the only prominent top-three entrant presented by its developer as open source, although its current public release does not yet appear to provide the weights needed for independent self-hosting.

For enterprises, that is not merely a licensing footnote: it affects deployment control, data residency, customization options, vendor dependence and potentially total cost of ownership.

Speaking before those final rankings were available, Guan said OpenArt’s preliminary voting was already surfacing cases it had not expected. “There are certain areas that surprised us,” she said. “Some models that we didn’t think would rank high in a particular use case — their advantages get more shown by this leaderboard.”

That, she argued, is part of the point of separating the tasks. “If you just compete on the overall, some models might not be able to be the best,” Guan said. “But they might be best in specific use cases or specific criteria.”

The final boards offer at least some evidence for that proposition, even if they do not overturn the market across the board. Seedance 2.5 looks less like a narrow specialist than a strong general video performer in OpenArt’s testing, but Wan’s first-place video-editing result and the changing image winners demonstrate why production teams may still want to route jobs among several models instead of standardizing blindly on a single one.

OpenArt isn’t the first or only to offer category-specific rankings

Leaderboard / benchmark

Who judges

How rankings are segmented

Evaluation approach

Key distinction

OpenArt Arena

Creative Expert Council + larger pool of creative “tastemakers”

Specific creative jobs: film, graphic design, e-commerce, motion design, video editing, lip sync, etc.

Curated task-specific prompt sets; blind pairwise comparisons; Bradley-Terry rankings

Most explicitly organized around individual production tasks across both image and video; integrated with a commercial creative platform

Contra Labs — Human Creativity Benchmark

Working professional creatives recruited from Contra’s network

Creative domains + workflow stages: landing pages, ad images, brand images, product video, etc., each tested across ideation, mockup and refinement

Blind pairwise comparisons + scalar ratings + qualitative feedback; Bradley-Terry/Elo; publishes detailed study methodology and datasets

More of an ongoing research program than a pure leaderboard; analyzes why performance changes across stages and where expert judgment converges or diverges

Arena.ai

Large-scale public/community voting

Image: commercial design, cinematic imagery, portraits, text rendering, 3D, art, etc. Video is segmented more by generation mode

Blind pairwise preferences with category rankings filtered from millions of real-world prompts; reports uncertainty/rank spread

Greatest emphasis on crowd scale and organic user prompts, its image rankings already overlap substantially with OpenArt’s task-specific concept

Artificial Analysis

Crowdsourced blind preference voters

Image use cases such as marketing, e-commerce, live-action film, UI/UX and creator content. Video is more function specific, including text-to-video, image-to-video, editing and audio variants

Curated prompt taxonomy; blind pairwise voting; Bradley-Terry Elo; also tracks price, speed and open-weight status

Broadest benchmarking/ procurement view, pairs quality with cost, latency and model openness; image categories are highly task-specific, while video segmentation is more modality-based than OpenArt’s creative-job boards

OpenArt is hardly the only firm offering rankings of AI model performance on creative tasks, however.

Contra Labs, for example, already runs practitioner-based evaluations of generative creative models as part of its Human Creativity Benchmark (HCB), which it introduced April 2026 as a way for creative practitioners to evaluate AI systems on professional creative work.

Rather than relying on generic technical metrics, Contra recruits working creatives from its network, gives models real-world briefs, blinds model identities during evaluation and uses pairwise comparisons aggregated with a Bradley-Terry model into Elo-style (chess) rankings.

Contra Labs' broader HCB framework spans domains including landing pages, desktop apps, ad images, brand images and product videos, and breaks work into ideation, mockup and refinement stages — a workflow-phase structure that differs from OpenArt Arena’s more granular rankings for specific creative tasks such as filmmaking, motion design, video editing and lip sync.

Contra’s benchmark is also designed to capture more than simple winner-versus-loser preferences. Evaluators rate outputs on prompt adherence, usability and visual appeal and provide qualitative explanations for their judgments. Contra explicitly distinguishes between “convergence,” where professionals tend to agree on issues such as legibility or technical correctness, and “divergence,” where disagreement reflects legitimate differences in taste and creative direction. Its research has repeatedly found that no single model dominates every stage of creative work — a conclusion that closely parallels OpenArt’s argument that the “best” model depends on the task at hand.

Beyond the core HCB, Contra Labs runs an ongoing research program of model battles, profiles and field notes across image generation, video, web design, typography, logo creation, style transfer and other professional creative tasks.

Recent studies have revisited the same domains as new models arrive or as Contra changes the evaluation structure, and the company publishes evaluator counts, prompt structures, judgment volumes and, in some cases, open datasets alongside its findings.

That makes Contra one of the clearest existing rivals to OpenArt Arena. The key distinction is less about methodology than how the products are organized and presented.

Contra currently operates more like a continuously updated research and benchmarking program, often analyzing where and why model performance changes across stages of a workflow.

OpenArt is launching Arena as a persistent public model-selection layer, with separate leaderboards for specific image and video jobs such as filmmaking, e-commerce, motion design, video editing and lip sync. In other words, practitioner judging, blind comparisons and workflow-aware evaluation are not unique to OpenArt; its differentiation lies more in packaging those ideas into a broad, continuously accessible set of task-specific leaderboards that it can also surface inside its own creative platform.

Arena.ai likewise already offers category-specific image rankings. It introduced category leaderboards after analyzing a large collection of real image-generation prompts, organizing them into areas including Product, Branding & Commercial Design; Photorealistic & Cinematic Imagery; Portraits; Text Rendering; 3D Imaging & Modeling; Art; and other domains. Arena.ai says a category board uses the same underlying evaluation methodology as its overall arena, filtered to prompts belonging to that domain.

The scale is in a different class from a curated expert panel. Arena.ai’s September 7 Product, Branding & Commercial Design text-to-image board alone reports more than 1.9 million votes across 78 models. Its image-edit arena similarly offers category views for commercial design, cinematic imagery, portraits, text rendering and other tasks.

As another prolific AI video creator, Kiri Margaros (@Kyrannio), creator of the automated video production studio No Spoon Studios told VentureBeat via direct message on X:

"Usually I look to Arena to evaluate what’s out there at present, and from there especially across text or image or ref-to-video, I pay closer attention to those metrics and models in my products when they test higher and debut across domains. If not using leaderboards, though, it actually comes down to just implementing the newer models as they come out into my product and then assessing how they fare from there with my current stack to see how it measures up. I think cost and speed are also tradeoffs though, as is overly aggressive moderation. If you have a super slow model that takes forever costs a lot and moderates everything you do, leaderboard positions don’t matter as much because it’s kinda unusable, though overall it’s a great and fast induction go quickly know which ones will be [state-of-the-art] SOTA, and with which strengths prior to fully diving in via API, if that tracks!! ... A lot of arena metrics for new models really guide my product roadmap specifically and I pay close attention to arena most of all there."

So “professionals judge the outputs” is not new, and neither is “different categories have different leaderboards.”

Where OpenArt does look meaningfully different is in the combination and organization of those ideas. Arena.ai’s image categories are largely prompt-domain clusters derived from large-scale community usage, meaning its specialized image rankings emerge by filtering votes from prompts classified into those domains.

OpenArt starts from the other direction: it first defines a professional use case, develops a dedicated test set and criteria for that board, and then asks its selected evaluator population to judge outputs against it.

Artificial Analysis is another close comparison — and on the image side, it narrows OpenArt’s differentiation further. Its Image Arena does not stop at one overall text-to-image ranking: Artificial Analysis says it ranks models across a taxonomy of real-world use cases including Marketing & Advertising, Retail & E-commerce, Live-Action Film, Animation & Gaming, Architecture & Real Estate, UI/UX Design, Social Media & Creator Content and others, alongside capability-specific cuts such as text rendering, layout and human anatomy.

It publishes corresponding subcategory views — including Commercial, People/Portraits and Text & Typography — and uses curated prompts tagged by both use case and capability. Outputs are judged through blind pairwise user voting, with ratings calculated using Bradley-Terry maximum-likelihood estimation and the prompt set refreshed monthly. That means OpenArt’s film and e-commerce image boards are not, by themselves, a new kind of task-specific image ranking.

The distinction is more meaningful in video. Artificial Analysis maintains separate Video Arena Elo pools for text-to-video, image-to-video and video editing, with additional splits for models that generate synchronized audio, and it publishes price and open-weight filters alongside the rankings.

But those are primarily divisions by generation modality or technical workflow, not the finer professional-use-case boards OpenArt is launching for film, motion design, video editing and lip sync. Artificial Analysis therefore already overlaps OpenArt substantially in image-task benchmarking and even directly competes on video editing, while OpenArt’s stronger differentiation appears to be the breadth of its creative-job-specific video rankings and its use of a selected expert council and creative-practitioner/tastemaker pool rather than Artificial Analysis’ broader crowdsourced preference votes.

Artificial Analysis also explicitly describes itself as independent and says it accepts no compensation from model providers for inclusion or favorable results — a structural distinction from OpenArt, which operates the commercial creative platform where many of the ranked models are available to users

For an enterprise creative lead, those methodological differences can change how a leaderboard should be used. A broad community-preference score can help identify models with general appeal, while a tightly scoped professional benchmark can be more useful for workflow routing.

Neither replaces internal testing on a company’s own brands, references, legal constraints and production standards. The practical value of Arena will therefore depend as much on how precisely OpenArt defines each board as on which model finishes first.

Put differently, OpenArt’s defensible differentiation is not that it invented expert evaluation or category segmentation. It is attempting to combine curated, profession-specific test sets, expert-led judging, a broader rater pool and a continuously accessible image-and-video leaderboard in one place.

It is also tying that benchmark more directly to model selection inside a commercial creative platform. That may make Arena unusually actionable for users who can move from a ranking to generating with the ranked model in the same environment, but it also introduces the most important governance question surrounding the launch.

The independence question

There is another distinction enterprise buyers should keep in view: OpenArt is not an independent academic benchmark organization. It is a commercial AI creation platform operating in the same market its leaderboard evaluates.

OpenArt itself says it aggregates more than 100 models, and its current model catalog includes systems such as Seedance, Veo, Kling, Wan, GPT Image, Nano Banana, Seedream and Grok Imagine.

That does not mean the launch results are biased, and the supplied rankings do not appear to include an OpenArt-branded foundation image or video model.

But OpenArt does sell access to many of the models being compared and operates products such as Director on top of them. Its video product explicitly offers users a range of underlying models inside the same workspace.

Arena will also be surfaced inside OpenArt’s own model-selection experience, Guan said, so the benchmark can influence which underlying model a customer chooses while remaining inside OpenArt’s platform.

That makes the relationship more complicated than a conventional third-party benchmark. OpenArt has a legitimate product reason to help users choose correctly — a customer who gets poor results from the wrong model may blame the platform — but it also has a commercial interest in becoming the place where customers both make that decision and execute it.

The appropriate conclusion is not that Arena’s results should be discounted, but that its methodology and governance deserve the same scrutiny enterprises would apply to any vendor-generated benchmark that could influence purchasing or workload-routing decisions.

That makes governance and reproducibility more important, not less. For launch, OpenArt has disclosed the statistical approach, the broad recruiting strategy, named members of its expert council and its plan to publish some prompts. But the company has not yet supplied the prompt count, completed judge count or total vote count that would let outsiders compare the scale and coverage of its launch evaluations with Contra’s published studies or Arena.ai’s massive community-vote datasets.

The comparison makes that omission more conspicuous because competitors have shown that those numbers can be disclosed without necessarily exposing an entire benchmark. Contra, for example, publishes its evaluator count, task structure and total pairwise judgments while still running repeat studies. Arena.ai exposes vote counts and confidence ranges directly alongside its rankings.

OpenArt’s planned 800-to-1,000-person tastemaker pool sounds substantial, but until OpenArt reports how many actually participated in a given board, how many judgments they produced and how many prompts they saw, readers cannot tell how much evidence sits underneath an individual launch ranking.

For enterprise teams, that should make OpenArt Arena an input-to-model selection rather than an oracle. The most useful benchmark is the one that resembles the organization’s actual work: product shots with exact packaging, campaign assets with typography, cinematic sequences with camera motion, or edits that must preserve a client’s existing footage.

OpenArt’s task-specific boards move in that direction, and the company also plans to include more objective enterprise-oriented model information such as pricing, content moderation and IP-protection characteristics. Guan said OpenArt also wants to expose price-performance tradeoffs rather than ranking quality without regard to budget.

The company says Arena will be public and “somewhat independent” from the rest of OpenArt, and Guan said OpenArt could eventually explore partnerships that allow other platforms to use or integrate the rankings. That is not part of the initial launch.

Where OpenArt Arena stands in the enterprise AI workflow

OpenArt Arena arrives at a useful moment because creative AI procurement is becoming a routing problem. The model that produces the best hero image may not be the one that renders typography cleanly; the best cinematic generator may not be the best editor; and the model with the highest aesthetic ceiling may be uneconomical at scale.

The first Arena results put a concrete model behind that argument. Seedance 2.5 is the standout of OpenArt’s launch video boards, taking the overall, film, motion-design and lip-sync rankings while missing the video-editing lead by a single point.

On images, leadership splits between GPT Image 2 for graphic design and Seedream 5.0 Pro for the film and e-commerce boards.

Those results are more useful as a map of where models appear strongest than as a proclamation of one permanent winner — especially in a market where, as Guan noted, another model can arrive a week later.

The launch results reinforce that point. The harder question now is whether OpenArt can make the benchmark itself transparent enough that creative teams trust the route it recommends.