
25 Sep 2026
Omneky launches TASTE BENCH to score AI ad creatives as ready-to-run
25 Sep 2026 (ET): Omneky introduced TASTE BENCH, an evaluation suite that tests how well image and video models turn creative briefs and brand assets into ready-to-run ads — with GPT Image 2.5 Sunburst posting the highest first-attempt ready-to-run rate at 54.2% across eight models.
SOFTWARE desk — brands are already sending the same brief to several image models and hoping one picture comes back as an ad they can run. A public ready-to-run score matters because a finished-looking image is not the same as a first try a reviewer would ship, and the best model in this edition still clears that bar on only 32 of 59 briefs.
What the release says the suite is for. Omneky, called the autonomous AI growth platform, introduced TASTE BENCH, an evaluation suite that tests how well image and video models turn creative briefs and brand assets into ready-to-run ads. The subhead says the suite assesses brand fit, creative execution, and whether an ad is ready to run. A creative brief, here, is the written request for the ad. Brand assets are the logo, the product, and the other files the brand already has. Those glosses are this desk’s. The scores below are for image ads. The release names video models in the suite’s job, and it does not print a video ready-to-run rate in this edition.
The image-ad score the release prints. For image ads, GPT Image 2.5 Sunburst has the highest ready-to-run rate at 54.2 percent, which is 32 of 59 briefs, with an average quality score of 7.38 out of 10. Rates across the eight models range from 22.0 percent to 54.2 percent. These are first-attempt results with Omneky’s own review step disabled. First-attempt means one generation, before that review step gets a chance to fix the picture. That gloss is this desk’s. 32 of 59 is a little more than half. The release does not say Sunburst also has the highest quality score. It says the highest ready-to-run rate, and it prints 7.38 as that model’s average.
How big this edition is. The current edition covers eight image models, 59 ad briefs, four brands, and five languages: English, Japanese, Arabic, Hindi, and Spanish. It counts 469 generated ads and 1,371 valid blind judge reviews. Blind, here, means the judge is not told which model made the image. That gloss is this desk’s. The release does not name the four brands. Sunburst is the only model this part of the wire names.
How the comparison is run. Models receive identical prompts and assets. A prompt is the instruction sent to the model. That gloss is this desk’s. A blind panel of AI judges from OpenAI, Anthropic, and Google scores outputs on eight quality dimensions, including brand fit, typography, and ad effectiveness. Typography is the type on the ad. Brand fit is whether the picture looks like that brand. Ad effectiveness, in this list, is a creative judgment, not a sales result. Those glosses are this desk’s. Eleven pass/fail checks cover exact copy, correct language, faithful logos and products, safe-zone placement, and fabricated claims or visual artifacts. Exact copy means the words match the request. A safe zone is the part of the frame an ad platform is not supposed to crop away. A fabricated claim is a line the picture states that the brief did not. A visual artifact is a broken or fake-looking part of the image. Those glosses are this desk’s. The release does not print the names of all eleven checks.
What ready-to-run means, and what it is not. An output counts as ready to run when it passes the applicable hard checks under the panel’s voting rules and a majority of judges would run it as-is. As-is means without another edit. Applicable hard checks are the checks that apply to that ad. That is the release’s phrase. It is not a claim that every check fires on every image. Results reflect automated creative judgments, not campaign conversions or return on ad spend. A conversion is the sale or signup the ad was meant to cause. Return on ad spend is the money back per dollar spent. Those glosses are this desk’s. 54.2 percent is not a revenue figure.
Where Omneky says the scores get used. The company’s AI Growth Agent chooses among image and video models from several labs for each ad. The about box says the agent researches strategy, creates ads and landing pages, launches campaigns, and allocates budget across Meta, Google, TikTok, ChatGPT, and more. It works inside Claude, ChatGPT, Grok, and Slack, routes each task to the best model, and learns from performance data on over $1 billion in ad spend. Founded in 2018 in San Francisco, Omneky serves over 6,000 customers. Several labs, best model, over $1 billion, and over 6,000 are the company’s words. This desk did not audit the spend or the customer count. The release points readers to the ads, briefs, judge explanations, scores, latency, and cost at www.omneky.com/tastebench. Latency is how long a generation took. That gloss is this desk’s.
Hikari Senju, as a quote, not as a result this desk measured. Senju, founder and CEO of Omneky, said taste is in the details: whether the typography feels right, whether the product looks authentic, and whether the idea comes through immediately. He said they built TASTE BENCH to make those judgments systematic, so they know which models to trust with their customers’ ads. That is his sentence on the wire. It is not a test this desk ran.
What the public explorer showed, kept apart from the wire’s unrounded figures. The page at omneky.com/tastebench says last updated Sep 24, 2026, the day before the 07:46 ET wire. The header this desk read counts eight leading image models, 59 real ad briefs across four brands and five languages, 469 ads generated, and 1,371 blind reviews. Those counts match the wire’s edition. The page says the same brand kit and the same brief, one shot each, scored blind by three judges. The all-briefs table, labeled 59 briefs in this view, prints Sunburst at 7.4 and 54 percent ready to run. 7.4 and 54 percent are that page’s labels. They are the rounded display of the wire’s 7.38 and 54.2 percent, not a second measurement. The sort this desk opened does not put Sunburst first. GPT Image 2.5 Flare is first on that table, at 7.4 and 42 percent ready. The ready-to-run labels on that view still run from 54 percent down to 22 percent, which matches the wire’s 22.0 percent floor. The page names all eight models. The wire names Sunburst only. The other names are in Sources. Filters on the page can shrink a view to 7 or 14 briefs. This filing’s 54.2 percent is the wire’s 59-brief figure, not one of those slices.
Where the explorer’s rule is worded differently. The page says an ad is ready to run only if it passes every check and most judges would run it as-is, and that a split vote counts as a fail. The wire says the applicable hard checks under the panel’s voting rules, and a majority of judges would run it as-is. Every check, and a split vote counts as a fail, are the explorer’s sentences. Applicable hard checks are the wire’s. Do not collapse them. The explorer also says three frontier AI models from three labs judge the ads, so no model grades its own homework, and that quality runs from 1, which it calls broken, to 10, which it calls agency-grade. The wire names judges from OpenAI, Anthropic, and Google, and eight quality dimensions. Three models, and those 1-to-10 labels, are the explorer’s. The page says it does not publish the recipe for how Omneky turns a brand into art direction, and that every model ran on that same recipe. That recipe line is not in the wire.
What the card shows. The card is the wire’s comparison image: Omneky’s TASTE BENCH product UI, with creatives from eight AI image models that received the same briefs and brand assets, each output beside creative quality scores, readiness assessments, and evaluation costs. That sentence is the wire’s caption on this image. A second image on the same release is a leaderboard, and its caption says Sunburst had the highest readiness rate in this evaluation at 54 percent. 54 percent there is that caption’s rounding of 54.2 percent. The card on this filing is the comparison UI, not that leaderboard. The frame has a purple margin. No date is printed on the card.
Plain English for the rest of the card. TASTE BENCH is a public test of whether an image model’s first ad is something a reviewer would actually run. In the edition Omneky published, the best ready-to-run rate is 54.2 percent, 32 briefs out of 59. The other models land between that and 22.0 percent. The judges are models from other labs, and they are not told who made the picture. The number is a creative grade. It is not how much money the ad made.
PRIMARY here: Omneky Inc.’s 25 Sep 2026 PR Newswire release 302890239, visible stamp Sep 25, 2026, 07:46 ET, datelined San Francisco — Tier A PRIMARY, the company’s own wire. STATUS PRIMARY. The desk label on the chip is CONFIRMED because that is the catalog word for a primary page we can stand on. The suite, the 54.2 percent and 32 of 59, the 7.38 out of 10, the 22.0 to 54.2 range, first-attempt with the review step disabled, eight image models, 59 briefs, four brands, five languages, 469 ads, 1,371 valid blind reviews, identical prompts and assets, judges from OpenAI, Anthropic, and Google, eight quality dimensions, eleven pass/fail checks, the ready-to-run rule, the line that these are not conversions or return on ad spend, the Growth Agent choosing among models, the Senju quote, the about box, the 2018 San Francisco founding, over 6,000 customers, and over $1 billion in ad spend are that wire’s. The explorer’s rounded 7.4 and 54 percent, the other model names, the Sep 24 last-updated line, the split-vote sentence, and the three-judge wording are the public page’s, read with this filing. They are not a second edition. The catalog chip is SOFTWARE, not SIGNAL.
RELATED
On 25 Sep 2026 Omneky introduced TASTE BENCH. The record is the company’s PR Newswire release, “Omneky Introduces TASTE BENCH for Evaluating AI Creative Quality.” The visible stamp is Sep 25, 2026, 07:46 ET, which is 7:46 a.m. Eastern and 11:46 a.m. UTC. The dateline is San Francisco, and it reads Sept. 25, 2026, with no second hour beside the city. A geo.region tag on the page says California. schema.org datePublished is 2026-09-25T07:46:00-04:00. That timestamp is 7:46 a.m. Eastern and 11:46 a.m. UTC, the same minute as the stamp. dateModified is 2026-09-25T07:46:55-04:00, fifty-five seconds later. This desk read the page as it stood. It did not diff those fifty-five seconds. The source line is Omneky Inc. A tracking pixel carries NewsItemId=PH56940 and Transmission_Id=202609250746PR_NEWS_USPR_____PH56940, DateId=20260925. That pixel is the wire’s own stamp echo. It is not a second announcement.
Sources
- PR Newswire — Omneky introduces TASTE BENCH, 25 Sep 2026
prnewswire.com
- Omneky — TASTE BENCH results explorer
omneky.com