This blog post was updated in July 2026 to reflect the most recent information, data, and insights available at the time of publication.
Let’s start with a statistic: According to our syndicated design research, nearly 60% of package redesigns fail to increase purchase preference compared to their predecessors. The brands that undertake these redesigns are not small or undisciplined, either; these redesigns very likely used standard end-of-process validation tools and met action standards. Nevertheless, when sampling redesigns that likely made it through these established processes, the failure rate is incredibly high.
We have tracked the sales data of more than 300 redesigns launched to market, cross-referencing their sales performance with our own consumer data—and the results showed the correlation we anticipated. Nearly every redesign that improved purchase preference also increased sales. Conversely, the designs that performed worse on this metric declined in sales. In other words, purchase preference (as measured by Designalytics) has a near-perfect correlation with sales outcomes.*

This chart showcases the results of legacy validation testing versus Designalytics. On the left are redesigns undoubtedly approved for launch using legacy end-of-process validation. On the right, Designalytics’ predictions for these same redesigns. Below each redesign are the sales results in the six months after the redesign was launched. As you can see, hundreds of millions could have been saved by using Designalytics' more predictive data.
And these are just 40 of 300+ redesign predictions we've validated using actual sales outcomes.
Why using legacy design research alone isn’t enough to assure success
Beyond the obvious issues with the failure rate, there’s something else at play. The consumer-packaged-goods (CPG) industry is laboring under the mistaken impression that design does not drive business outcomes because they have seen no clear connection between the validation scores a redesign receives and the sales of that product once the redesign launches. Because brands and marketing mix analysts couldn’t see the impact, they assumed there wasn’t any to speak of.
The truth is: Design doesn’t have an impact problem. It has a measurement problem.
Or, rather, it did have one. Designalytics’ design measurement tools can be used to test pre-market designs with higher than 90% predictive reliability of whether the redesign will outperform its predecessor in sales. Our unmatched data quality, massive sample sizes, innovative metrics, advanced exercise design, and much more have completely changed quantitative package design testing.
This does beg the question, though: Why do so many redesigns that leverage traditional research tools fail?
There can be many reasons. Here are a few:
Testing too late in the process
Historically, many large brands have chosen not to conduct robust research that would help guide creative strategy at the beginning of the design process (an unfortunate consequence, perhaps, of the mistaken belief that design is not impactful). Instead, they’ve focused on intensive research, such as shelf tests intended to replicate store environments, at the end of the process to validate their chosen design route. This is much too late to provide vital direction to the creative team.
The result of this approach is less informed creative strategies, an increase in subjective decision-making, and limited opportunities for creative exploration and refinement that could've created higher-performing designs. Importantly, though, the trend has shifted: Brands have seen the value of doing research earlier in the creative process.
Parity-or-better action standards
In general, the accepted measure of success for a traditional validation design test is “parity or better.” Now, if you asked a random collection of brand managers (or people in general) what “parity” means, most would probably say “equal” or “as good as.” For good reason, too: that’s basically the dictionary definition.
In traditional validation testing, however, parity is a statistical term with a divergent meaning: that the sample size is too small to have confidence that the new design is better, worse, or exactly the same. In other words, parity amounts to a very expensive, anticlimactic shrug.
To make matters worse, parity is the outcome for a vast majority of design tests because lower sample sizes (i.e., 100-150) are often used. Given the confluence of these factors—the subpar success rate and the frequency of parity results—the bar for design performance metrics has been lowered to accommodate the high incidence of parity results. As a consequence, brands are given the “green light” more often… but often at a significant cost.
Brands rarely question this entrenched idea because they think it means design doesn’t have a significant impact.
As our data consistently demonstrates, it has a major impact. In fact, our design research arrives at a parity result for less than a fifth of our redesign tests—which means we deliver a decisive and predictive outcome around 80% of the time. That kind of clarity leads to a higher degree of confidence; Brands can see clearly which design works better and why, without a post-hoc dive into the numbers to justify their decision.
Focusing on the wrong measures
Again, this doesn’t mean measures like findability shouldn’t be considered; it simply means that some metrics are demonstrably more important when it comes to in-market outcomes. Keeping this in mind allows for more effective designs when trade-offs are necessary.
Failure to evaluate the current design’s performance
There’s a reason the adage isn’t “If it ain’t broke, let’s invest immensely in an uncertain alternative.” The truth is that some package designs are performing well already. Assessing your package design regularly arms your organization with the insights to know when a redesign is warranted and, if so, what changes will likely drive growth. Otherwise, decisions can be based on fear, conjecture, or just opinion.
When it comes to package design, the ill-fated Tropicana redesign of 2009 has become a cautionary tale. The esteemed OJ brand’s revamp was wildly misguided, and consumers roundly rejected it. In a matter of months, the brand had lost $30 million in sales.
Yet, Tropicana redesigned its orange juice bottle in 2024 and, once again, consumers revolted. The issue was different from its 2009 predecessor—this time, the altered bottle shape is primarily what’s driving consumers away, rather than the label design—but the impact is the same. Sales plummeted for Tropicana’s orange juice products, and they lost 4 percentage points of market share to one of its top competitors, Simply Orange.

There were so many opportunities for this to go in a better direction. For one, a baseline assessment of the design would have likely shown the brand the folly of messing with distinctive brand assets and design elements that were working. Testing design concepts with consumers earlier in the process would have illuminated red flags and likely led to successful refinements. This latest design no doubt passed legacy validation testing, which shows just how limited it is. Having access to more predictive data at the end of the process would have likely prevented this false positive, saving millions of dollars. Each redesign has its own story of success or failure. Designalytics offers brands the opportunity to see potential failures before they arise, leverage objective consumer insights to maximize successes, and actually see the impact your design can have on the growth of your brand.
Clarity. With the design measurement tools available to brands today, it’s very easy to attain. If you’re considering a redesign, we invite you to request a demo so that you can see firsthand the difference the right data at the right time can make.
*This correlation is binary. Designalytics reliably predicts whether sales will increase or decrease, but not by how much.


