TV Commercial Testing: How to Measure Effectiveness Before Air

TV Commercial Testing: How to Measure Effectiveness Before Air

TV Commercial Testing: How to Measure Effectiveness Before Air

TV commercial testing, also called copytesting, is the process of measuring how a television ad will resonate with audiences before it airs. It typically evaluates recall, persuasion, emotional response, and brand linkage using animatics, photomatics, or finished cuts, so weak creative can be fixed or replaced before broadcast media spend is committed.

TV Commercial Testing

Tag

Research

Date

Read Time

8 Min

Content

Senior Growth Marketer

Summary:

  • TV commercial testing, which is known as copytesting, is a method of finding out how an audience reacts to an advertisement before it is broadcast.

  • This is important since broadcast budgets are committed at the beginning and cannot be recovered once a poor commercial has been aired.

  • The main methods involve copytesting, animatic and photomatic testing, and the measurement of emotional responses in the areas of recall, persuasion, and brand linkage.

  • Make the testing early so as to allow the creative process to be altered, and always compare the results with benchmarks that are appropriate to the stage.


Television still commands some of the largest single line items in a marketing budget, and unlike a digital placement, a broadcast buy cannot be paused halfway through a flight without losing most of its value. That makes the quality of the creative itself the biggest variable in the outcome. Nielsen's analysis of nearly 500 campaigns found that creative contributed 47% of a campaign's sales impact, more than reach, brand, or targeting. In other words, the spot you choose to run matters more than almost anything else you decide during planning.

Pre-air testing is a method that helps teams cut down on the uncertainty when making that decision. This guide explains what TV commercial testing is, how it is different from in-market measurement, which methods and metrics to use, and also how to carry out a test that actually results in a change to the programme that goes to air.

What TV Commercial Testing Means

TV commercial testing involves measuring the audience's reaction to a television advertisement before it is aired. Ever since decades ago the industry has referred to this as copytesting and the main objective has remained the same: to discover whether or not a commercial is likely to be successful before spending money to show it to millions of people.

You don't need the film to be completed in order to carry out testing. A commercial can be assessed at almost any stage of production, whether it's a rough animatic based on storyboard frames, a photomatic put together from still images, or the final graded cut. The thing that differs at those various stages is the degree of precision, not the usefulness. Tests carried out in the early stages show you whether the idea succeeds in making an impact, while those done at a later stage show you whether the execution meets that impact.

The output of a test is a decision, not a report: run the spot as is, revise it before air, or drop it and back a different execution. That is why TV commercial testing sits so close to broader creative effectiveness work. Both separate creative that will earn its media investment from creative that will quietly waste it.

Why Pre-Air Validation Matters for Broadcast Spend

Television produces an unusual imbalance in costs: the pre-broadcast testing involves a small and fixed amount of research expenditure, while the media company's investment in the spot is large and mostly irreversible. After a booking has been made, any corrections have to either withdraw the spot—which wastes the purchase—or carry on with it, which wastes the impressions.

That asymmetry gets worse when the creative is weak. Kantar's validation work on its Link ad testing database shows how wide the gap is: an ad scoring highly on short-term sales likelihood typically generates 33% more sales than an average ad, while a low-scoring ad generates 27% fewer. The same media plan, behind two different spots, can produce results that differ by roughly 60 percentage points.

There is also a reputational aspect to it. If a digital advertisement makes a mistake about its tone it will reach a particular group, whereas a television commercial addresses a large audience all at once, usually during moments of high attention such as live sports, and a misjudged tone can turn into a news story within a few hours.

Then there is the budget conversation. McKinsey has found that companies can unlock savings of 10 to 20 percent by eliminating inefficient spend and redirecting it toward higher-performing activity. Killing one underperforming spot before it airs is one of the clearest versions of that discipline, and it is why structured TV ad testing has become a standard gate in most brand teams' production timelines.

TV Commercial Testing vs In-Market Measurement

These two things are often confused, and the distinction is important.

Pre-air testing predicts. It asks a relevant sample of viewers to watch a commercial in a controlled setting, measures how they respond, and compares those responses to norms from similar ads. It happens before any media is bought.

In-market measurement confirms. Brand lift studies, sales modelling, and attribution describe what actually happened after a spot ran. They are more definitive, but they need real airtime and real spend to generate a signal, and by the time results arrive the money is gone.

The relationship is sequential rather than competitive. Use pre-air testing to select and sharpen the creative, and in-market measurement to confirm impact and build norms for the next round. Teams that only do the second half are learning at full media rates. It also helps to know what each stage explains: pre-air work is strong on diagnostics like where attention drops and whether the brand registers, while in-market work is better at commercial outcomes. Understanding the relationship between attention and recall keeps teams from over-reading a single pre-air number.

TV Commercial Testing Methods

There is no single method called "TV commercial testing." It is a family of approaches that differ by production stage, by what they measure, and by how fast they return results. Most brands use two or three together, and the mix has broadened as ad creative testing platforms have added behavioural and automated measures alongside traditional survey questions.

Copytesting

Copytesting is the general term used for assessing a creative presentation against predefined criteria before it is aired. A sample of viewers sees the advertisement and then completes a structured set of questions covering recall, understanding of the message, level of persuasion, likeability and association with the brand, the results being compared with a normative database so that the score has meaning.

It has been the usual approach in television advertising since the 1960s, and this is indeed both its advantage and its drawback. Although the normative databases are extensive, the method depends on what people can consciously recall following an advertisement. Contemporary ad testing programmes typically combine the survey battery with behavioural measurements rather than replacing the survey battery.

Animatic and Photomatic Testing

Animatics are rough moving versions of a storyboard, usually with scratch voiceover and simple transitions. Photomatics use still photography or stock footage edited to the intended timing. Both exist to test the idea before committing to production costs that can run into six or seven figures.

The key point in this situation is benchmarking. An animatic will generally do poorly compared to a completed film in terms of likeability and production appeal simply because it appears to be unfinished, which is why animatic scores have to be measured against animatic norms. If you get this wrong, then you will end up ruining good ideas for the wrong reason.

Early-stage work is where low-cost creative tests that predict ad performance earn their keep, because changing direction is still cheap. Many teams pair this with concept testing so the idea and its execution are validated in the same window.

Emotional Response Measurement

The measurement of emotional response monitors the way that viewers react from moment to moment throughout a commercial rather than taking a single score at the end; instead of asking whether or not a person liked the ad, it records where engagement reached its peak, where it decreased, and whether the emotional climax occurred at the same time as the brand moment.

Emotion is not decoration in television advertising. Harvard Business Review research found that fully connected customers are 52% more valuable on average than merely satisfied ones, which makes the emotional arc of a 30-second spot a commercial variable rather than an aesthetic one.

Measurement is typically done through facial emotion AI, which reads expression frame by frame through a standard webcam. This primer on facial coding in marketing covers how the signals are captured and what they can and cannot tell you. The method works best alongside recall and persuasion measures, not as a substitute for them.

What to Measure Before a Commercial Airs

A pre-air test is only as good as the metrics it collects. Four groups matter most.

Brand recall and ad recall.

Can viewers remember the commercial a day later, and can they name the brand unaided? These baseline memorability measures are the ones most closely tied to media weight decisions.

Persuasion and purchase intent.

Does the spot shift how people feel about buying? These measures are directional rather than predictive of exact sales, but a commercial that moves nobody is unlikely to move a market.

Brand linkage.

The metric most often neglected and most often responsible for waste. Viewers may love the story, remember the joke, and attribute it to your competitor. Strong linkage means the emotion attaches to the right name. Weak linkage means you have funded category advertising.

Message clarity.

If the intended takeaway does not survive a single viewing, nothing else in the test matters much. This is where message testing belongs, ideally before final production.

Reading any of these in isolation is risky, and category context matters more than industry averages. Recent attention benchmarks from the state of creative testing show how much variation sits between formats, which is why generic norms mislead.

How to Run a Pre-Air TV Commercial Test

A workable pre-air test comes down to three decisions.

Step 1: Choose the production stage you are testing at.

Work backwards from the air date. With eight weeks and budget flexibility, test at animatic stage so findings can change the shoot. If the spot is already cut, test the finished film, but be honest that your options are limited to edit-level changes. Testing late and calling it validation is a common and expensive habit.

Step 2: Build a sample that reflects who will actually see the ad.

A general-population panel gives you a clean-looking result that describes nobody in particular. Screen for category buyers and account for where the spot will run. Nielsen's Q1 2026 Ad Supported Gauge shows that ad-supported viewing accounts for nearly 73% of total TV consumption, with streaming holding a record 46.6% share of it, so a commercial destined for a mixed linear and streaming plan needs a sample reflecting both viewing behaviours.

Step 3: Benchmark, then decide.

Compare results against stage-appropriate and category-appropriate norms, then commit to one of three outcomes: run it, revise it, or replace it. The upside of acting is substantial, with Kantar's analysis suggesting that lifting an ad's creative quality from average to great can increase return on investment by around 30%.

Teams that get the most from this treat it as a repeatable gate rather than a one-off study. A consistent approach to ai creative testing across campaigns builds an internal normative database, which makes every later decision faster and better informed.

Common Mistakes in TV Commercial Testing

Comparing rough creative against finished norms. Covered above, but worth repeating because it remains the single most common methodological error. Match the benchmark to the asset.

Relying on one metric. Recall alone tells you a commercial was noticed, not that it did anything useful. Persuasion alone tells you people said the right thing without proving they will remember it. Brand linkage tells you attribution without telling you appeal. The picture only forms when they are read together, which is the core argument for testing TV and display ads with emotion AI rather than survey questions alone.

Testing too late to change anything. If results arrive a week before air, the test is a post-rationalisation exercise. Build the testing window into the production schedule from the start.

Ignoring wear-out across the flight. A spot that tests well can still lose effectiveness after heavy rotation. Planning for creative fatigue means testing enough variation to refresh the flight rather than running one execution until the numbers sag.

Measuring Emotional and Attention Signals Before a Spot Airs

Traditional copytesting asks viewers what they thought. Behavioural measurement observes what they did while watching, which matters because most of the response to a television commercial is not consciously accessible. Three signals are particularly relevant to pre-air validation:

  • Facial expression, showing the emotional arc second by second and whether the peak lands near the brand moment or well away from it.

  • Visual attention, captured through eye gaze tracking, showing whether viewers actually look at the pack shot, logo, or on-screen offer. This explainer on eye gaze tracking and visual attention covers how webcam-based gaze data is collected and read.

  • Sustained attention, where attention measurement pinpoints the frames at which engagement drops, so an edit can be tightened rather than guessed at.

Together these turn a single end-of-ad score into a timeline a creative director can act on. Instead of "the ad scored below norm," you get "attention holds through the first twelve seconds, drops during the product demonstration, and the logo appears three seconds after the emotional peak has passed."

For teams screening several cuts or competing storylines before locking a spot, Predictive Creative AI offers a faster first pass, scoring multiple executions before human testing narrows the field. Decode's creative insights platform combines these signals in one workflow, and this guide to AI creative testing for ad performance explains how the automated and human-response layers fit together. The same signals were used to rank the top Super Bowl ads of 2026 by emotional response, which shows the method applied to television advertising at its most scrutinised.

Frequently Asked Questions

1. How much does TV commercial testing cost?

Cost varies with sample size, number of markets, and whether you add behavioural measures to a survey-only design. Automated platform tests are a small fraction of full-service custom research, and both are a fraction of a national broadcast flight. The relevant comparison is test cost against the media at risk, not against zero.

2. Can you test a TV commercial before it is fully produced?

Yes, and it is usually the better choice. Animatics and photomatics allow testing at script or storyboard stage, when changes are still cheap. The trade-off is that rough executions must be benchmarked against rough-execution norms.

3. What is the difference between copytesting and brand lift studies?

Copytesting happens before air and predicts likely performance. Brand lift studies happen during or after a campaign and measure actual shifts in awareness, consideration, or favourability among exposed audiences. Copytesting informs which spot to run, and brand lift tells you what running it achieved.

4. How long does TV commercial testing take before air date?

Automated tests can return results in a day or two, while custom qualitative-plus-quantitative programmes take two to four weeks. Work backwards from the air date and leave room for at least one revision cycle, otherwise the result cannot influence the creative.

5. What sample size do you need for a reliable copytest?

Most normative databases are built on samples of roughly 150 to 300 target-audience respondents per execution. Smaller samples work for directional early-stage reads, but they widen confidence intervals to the point where close comparisons between cuts become unreliable.

6. Is recall or persuasion a better predictor of TV ad effectiveness?

Neither works well alone. Recall predicts memorability, persuasion predicts attitude shift, and brand linkage determines whether either benefits your brand. Validated models combine short-term sales measures with longer-term brand equity measures rather than relying on one headline number.

7. Does TV commercial testing work the same way for CTV and streaming ads?

The metrics are largely the same, but viewing context differs. Streaming environments involve shorter pods, different screen sizes, and often more attentive viewing. Benchmarks should be format-specific, and skippable or interactive formats need measures that linear norms do not cover.

Test the Spot Before You Buy the Airtime

The case for pre-air validation is not complicated. Creative is the largest controllable driver of advertising performance, broadcast money is committed upfront, and the gap between a strong and weak execution on the same media plan runs into tens of percentage points. Testing before air is the cheapest point at which a bad decision can still be reversed.

Decode by Entropik brings emotion, attention, and eye-tracking signals into pre-air commercial validation, so teams can see where a spot holds viewers and where it loses them before the first flight is booked.


From Emotion to Action, With Insights That Speak Your Language.

Start turning customer signals into smarter decisions.

From Emotion to Action, With Insights That Speak Your Language.

Start turning customer signals into smarter decisions.

From Emotion to Action, With Insights That Speak Your Language.

Start turning customer signals into smarter decisions.