More than 33 exabytes of data cross the internet every day, and streaming video makes up around 54% of everything flowing downstream, according to Sandvine's 2024 Global Internet Phenomena Report. Almost every one of those videos passed through a codec before it reached a screen. Codecs are what make that possible. An hour of uncompressed 1080p video is about 336 gigabytes. The same hour streams to a phone in around two gigabytes. Nothing else in a video product does that much work, and the codec doing it decides how much a company pays for bandwidth, how good the picture looks, how warm a phone gets, and whether playback works at all on a five-year-old Android device.
So "which codec should we use?" deserves an answer based on evidence rather than instinct. Codec vendors publish impressive numbers, and those are usually true, for their test content and on their hardware. Whether they hold for a different kind of video is a separate question, and answering it is what codec benchmarking means.
In this blog post we will look at what a codec actually does, the three consequences that make codec choice a business decision as well as a technical one, how video quality gets measured when "it looks fine to me" is not good enough, and the five principles that separate a benchmark you can trust from one that quietly misleads you.
Every video users watch has been squeezed by an encoder first. How well that squeeze works is what codec benchmarking measures.
TL;DR
30-second summary
What does it actually mean to benchmark video codec performance, and how do you do it without the result quietly misleading you?
- Codec choice is a business decision with three separate consequences, not just a technical one. It sets the bandwidth bill a company pays, since better compression reduces CDN data costs almost proportionally. It drains batteries and heats phones when a device lacks hardware decoding support and falls back to software. And it decides who can watch at all, since newer codecs like AV1 and HEVC don't reach every device.
- Codec performance pulls in three directions at once, and no single score captures all of them. Compression efficiency measures how little data is needed for a given quality level. Computational cost covers the processing power and energy spent encoding and decoding. Compatibility measures how much of an audience can actually play the result. AV1 wins decisively on compression, loses on computational cost, and loses on compatibility, which is why a benchmark measuring only one dimension is nearly useless.
- Three objective metrics do most of the work in measuring video quality, and each has a distinct character. PSNR compares pixels directly and is fast but doesn't match human perception well. SSIM looks at structural patterns and is good at catching blocky artifacts. VMAF, developed by Netflix, combines multiple measurements trained against real human viewer scores and comes closest to predicting what a person would actually say.
- Five principles separate a trustworthy benchmark from one that quietly misleads. Compare codecs at equal measured quality rather than equal settings, since the same numeric setting means something different across codecs. Test content that resembles your own product. Measure across a range of bitrates rather than a single point, since quality flattens out and that flattening point is often the most valuable finding. Change one variable at a time. And verify the measurement itself, since a single frame of misalignment can shift a PSNR score enough to reverse a conclusion.
- Lab benchmarking on one machine only answers part of the question. It can tell you which codec compresses your content most efficiently, but not what users actually experience on real decoders, real networks, and real batteries. Getting that answer requires moving to device labs and controlled network conditions. TestDevLab's own conferencing platform benchmarking work reached 360 tests producing 2,304 video files across 14 devices and three bandwidth conditions.
Bottom line: A codec decides what a business pays to deliver video, how much battery and heat playback costs users, and how many of them can watch at all. Those consequences are why the decision deserves measurement rather than a vendor's claim, and measuring it well is mostly a matter of discipline — equal quality comparisons, representative content, a full bitrate range, one variable at a time, and a verified measurement.
What is a video codec?
A codec is really two programs working as a pair. The encoder squeezes video down to a smaller size, the decoder unpacks it again for playback. Put the halves together and you get the name: co-dec. The four you will meet most often are H.264 (also called AVC), H.265 (also called HEVC), VP9 and AV1.
Here is a distinction that trips up a lot of people at first. A codec is a specification, a written standard. The thing you actually run is an implementation of it, and there are usually several. H.264 is one standard - x264 is the popular software encoder for it. AV1 is another standard, SVT-AV1 is one encoder for it. Your phone and GPU have hardware encoders built in too.
Why video has to be compressed at all
Uncompressed video is absurdly large. A single 1080p frame holds about two million pixels, which works out to roughly 3 MB of raw data. At 30 frames per second that is 93 MB every second, or about 746 megabits per second flowing continuously, roughly 336 gigabytes for a single hour. No home internet connection can carry that. A typical 1080p stream arrives at around 5 megabits per second instead. The codec's job is to bridge those two numbers - throwing away roughly 149 parts in 150 while keeping the video looking essentially unchanged.
It manages this mostly by noticing repetition. Most of any frame looks a lot like the frame before, so instead of storing every frame in full, the encoder stores the differences. It also discards detail that human eyes are bad at noticing, such as fine gradations of colour. Better codecs are simply better at finding that redundancy and better at judging what nobody will miss.
The codecs you will hear about
Four names dominate, and they arrived in roughly this order:
| Codec | Introduced | Reputation |
|---|---|---|
| H.264 / AVC | 2003 | The universal safe choice. Plays almost everywhere. |
| H.265 / HEVC | 2013 | Much better compression, complicated patent licensing. |
| VP9 | 2013 | Google's royalty-free answer, used heavily by YouTube. |
| AV1 | 2018 | Royalty-free, the best compression available, resource intensive to encode |
If you want to see how these codecs compare, we suggest you read: Video Quality Comparison for Various Codecs at Different Resolutions and Bitrates (Part 3)
Why codec performance matters
Three separate consequences, each landing on someone different.
It sets the bandwidth bill
Content delivery networks, the services that store copies of your video around the world and serve it to nearby viewers, charge for data leaving their servers. Reduce the data needed per video and that bill falls almost proportionally.
The savings are large enough to matter. In a comparison run by the Streaming Learning Center, the same football clip was encoded with each codec and tuned so that all four looked equally good to a quality model. H.264 needed 9.0 MB, HEVC 7.1 MB, VP9 6.8 MB, and AV1 just 4.1 MB. Same apparent quality, 54% less data.
But better compression is not free. Squeezing harder takes more computing power, and the same study estimated AV1 encoding at roughly four times the cost of H.264. That produces a genuine break-even calculation. Namely, a video has to be watched somewhere between 950 and 4,000 hours in total before the delivery savings repay the extra encoding. Popular content clears that bar easily, but a large library of rarely-watched videos may never clear it at all.
It drains batteries and heats phones
Unpacking video takes work too, and this is where users feel codec choice directly.
Modern phones and laptops contain dedicated hardware for decoding common codecs, which is fast and uses very little power. If a device does not have the right hardware for a codec, it has to do the work using the main processor instead. This is called falling back to software, and it costs far more power. The battery drains faster, the phone gets warm, and once it gets too warm the processor deliberately slows down to cool off. At that point the video starts to stutter and skip frames.
Curious how your app performs under real playback conditions?
Battery drain and thermal throttling don't show up in a spec sheet, they show up on a real device after fifteen minutes of use. Our battery and data usage testing service measures exactly that, on real hardware rather than estimates.
Playback cost is also the part most often missing from codec comparisons, including published ones. A codec can compress brilliantly and still be the wrong choice for a mobile app, simply because playing it back costs too much.
It decides who can watch at all
Compatibility is not something you balance. A device can either play the video, or it cannot. A survey of 1.14 million browser sessions in early 2026 found H.264 and VP9 supported almost everywhere, at 99.94% and 99.99%, while AV1 reached about 91.5% and HEVC about 85.1%. Those numbers are a rough guide. The survey records what browsers report rather than what devices can really do. The ranking is what matters, and it matches what the rest of the industry sees.
The practical result is that newer codecs are almost never deployed alone. They ship alongside an older fallback, so that whatever a viewer's device supports, something plays. That means encoding and testing several versions of every video instead of one, which is a real cost that arrives attached to the compression saving.
What "performance" actually means here
There is no single score, and expecting one is the first mistake. Codec performance pulls in three directions at once.
Compression efficiency is how little data is needed to reach a given level of visual quality. Computational cost covers the processing power, time and energy consumed at both ends - encoding and decoding. Compatibility is how much of your audience can play the result smoothly.
Newer codecs win decisively on the first, usually lose on the second, and often lose on the third. AV1 is the clearest example: unbeatable compression, expensive encoding, incomplete device support. So a benchmark that measures only compression efficiency will recommend AV1 every time, which makes it useless as a benchmark because it was never capable of producing a different answer.
The useful output is a statement like: this codec saves 40% of our bandwidth, costs four times as much to encode, and cannot reach 8% of our users without a fallback. That is something a team can actually decide on.

Compression savings show up on the bandwidth bill. The cost of achieving them shows up in compute and in the user's battery.
How video quality gets measured
"It looks fine to me" does not scale. Comparing four codecs at four quality settings means sixteen files, and human eyes are inconsistent, slow, and easily influenced by knowing which file is which. So the industry uses objective metrics. Algorithms that compare the compressed video against the original and produce a number.
Three of them do most of the work, and each has a distinct character.
PSNR compares the two videos (compressed version against the original) pixel by pixel and reports the difference in decibels. It is fast, universally available, and useful for spotting when something has changed. Its weakness is that it does not match human perception especially well. Two files can score identically and look noticeably different, particularly where blurring is involved.
SSIM looks at brightness, contrast and structural patterns rather than individual pixels, producing a score between 0 and 1. Above about 0.97, degradation is effectively invisible. It is good at catching the blocky artifacts people actually complain about.
VMAF was developed by Netflix and has become the default for streaming work. It combines several different measurements using a model trained on scores from real human viewers, and reports a result on a scale from 0 to 100. Of the three, it comes closest to predicting what a person would say.
Two habits make these metrics trustworthy. First, use at least two of them, because when they disagree sharply, that gap is telling you something, usually that fine detail or grain has been lost, and it is worth opening the frames to see what. Second, remember what these scores actually are: software predicting what a person would probably say, and they are usually right. When one surprises you, go and look at the video rather than trusting the number.
We explore how these three metrics behave, including cases where they contradict each other, in our article Full-Reference Quality Metrics: VMAF, PSNR and SSIM. For how content type changes the picture, Temporal and Spatial Features for Video Quality Assessment is a useful companion.
Five principles of a fair benchmark
The tools matter far less than the method. These five principles make the difference between results you can use and results that only look right.
Compare at equal quality, not equal settings
Every encoder has a quality dial, and the numbers on those dials mean completely different things between codecs. Setting all four codecs to "23" and comparing the resulting file sizes feels like a fair test but in reality it is not. It’s like comparing a size 8 shoe in the UK against a size 8 in the US.
A fair comparison tunes each codec until all the outputs reach the same measured quality, then compares what each one needed to get there. The industry has a standard measure for exactly this, called BD-Rate, which reports the average data saving between two codecs at matched quality.
Test content that resembles yours
Compression performance depends heavily on what is in the video. A football match, with constant motion and grass texture, is a completely different problem from a person talking to a webcam against a plain wall. Gains measured on one do not transfer to the other.
Test two or three short clips that reflect what your product genuinely carries. This single choice affects results more than most people expect, and it is the usual reason a benchmark you read about gives different numbers when you run it yourself.
Test a range, not a single point
One measurement tells you almost nothing, because the relationship between data and quality is a curve rather than a line. Quality climbs steeply at low bitrates, then flattens out.
Finding where it flattens is often the most valuable result of the whole exercise. In our own testing across resolutions and bitrates, every codec was already close to its best possible quality by around 2,500 kbps, and spending more bought almost nothing anyone could see. H.265 reached its best 1080p scores at just 500 kbps. That flattening point marks exactly where a product starts paying for data nobody perceives, and testing at a single setting will never reveal it.

AV1 needs less than half the data H.264 does for the same measured quality, but takes around four times as much computing power to produce it.
Change one thing at a time
Encoders have dozens of settings, and several of them shift results as much as the codec itself does. The most dramatic is the effort setting, which controls how hard the encoder searches for savings. In one published test, moving between the slowest and fastest effort levels changed encoding speed by a factor of 124 while barely moving the quality score.
If two things change between runs, the result tells you nothing about either. Fix every setting, record what you fixed, and change one variable per comparison.
Verify the measurement itself
This is the principle beginners skip, and it can invalidate everything.
Quality metrics work by comparing the first frame of one file against the first frame of the other, then the second against the second, and so on. If the two files are misaligned by even a single frame, the comparison is between the wrong pairs, and the resulting score looks perfectly normal. A misalignment of even one frame can shift a PSNR score by several decibels, easily enough to reverse a conclusion.
Where it gets harder
Everything we mentioned above can be done on one machine, and it answers a useful question: which codec compresses our kind of content most efficiently.
What it cannot tell you is what users will actually experience. That depends on real decoders on real devices, real network conditions, and real batteries - and none of those exist on a test machine. Getting to that stage means moving from a laptop to device labs and controlled networks. We cover that pipeline in How We Automate Audio and Video Quality Testing, and How to Produce Research-Grade Benchmarking Data on Conferencing Platform Performance shows the scale a genuinely defensible comparison reaches: 360 tests producing 2,304 video files across 14 devices, three internet providers and three bandwidth conditions.
The principles do not change, there is simply far more of it, and it has to be automated.
Conclusion
A codec is the reason an hour of video takes two gigabytes instead of 336. Because it does that much work, the choice of codec decides three things simultaneously: what a business pays to deliver video, how much battery and heat playback costs its users, and how many of them can watch at all.
Those consequences are why the decision deserves measurement rather than a vendor's claim. And measuring it well is mostly a matter of discipline rather than expertise. Compare at equal quality instead of equal settings, test content that resembles your own, measure across a range so you can see where quality stops improving, change one variable at a time, and confirm your measurement is comparing what you believe it is.
Follow those and the numbers will point somewhere worth going. Skip them and you will still get numbers, equally confident, and pointing at the wrong answer.
FAQ
Most common questions
What is a video codec and how is it different from a standard?
A codec is really two programs working as a pair. The encoder squeezes video down to a smaller size, and the decoder unpacks it again for playback. A codec like H.264 or AV1 is a written specification, while the thing actually running is an implementation of it, and there are usually several implementations per standard, such as x264 for H.264 or SVT-AV1 for AV1. Phones and GPUs also typically have hardware encoders and decoders built in for the most common standards.
Why does codec choice affect more than just video quality?
Codec choice carries three separate consequences that land on different people. It sets the bandwidth bill a business pays, since content delivery networks charge for data leaving their servers and better compression reduces that cost almost proportionally. It affects battery life and heat on the playback side, since a device without hardware decoding support for a given codec has to fall back to software decoding, which costs far more power. And it determines compatibility, since newer codecs like AV1 and HEVC don't reach every device, which is why most products ship a modern codec alongside an older fallback rather than choosing just one.
What is the difference between PSNR, SSIM, and VMAF for measuring video quality?
PSNR compares two videos pixel by pixel and reports the difference in decibels. It's fast and universally available, but doesn't match human perception especially well, since two files can score identically and still look noticeably different. SSIM looks at brightness, contrast, and structural patterns rather than individual pixels, producing a score between 0 and 1, and is good at catching the blocky artifacts people actually complain about. VMAF, developed by Netflix, combines several measurements using a model trained on real human viewer scores and comes closest of the three to predicting what a person would actually say about the video's quality.
Why should codecs be compared at equal quality rather than equal settings?
Every encoder has a quality dial, but the numbers on those dials mean completely different things between codecs, so setting all codecs to the same numeric value and comparing file sizes produces a misleading result. A fair comparison instead tunes each codec until all outputs reach the same measured quality level, then compares what each one needed in data to get there. The industry standard for this is BD-Rate, which reports the average data saving between two codecs at matched quality, giving a genuinely comparable result rather than one that depends on arbitrary setting conventions.
Why is testing across a range of bitrates more useful than testing at a single setting?
The relationship between data and quality is a curve, not a straight line. Quality climbs steeply at low bitrates and then flattens out, and finding exactly where it flattens is often the most valuable result of the entire benchmarking exercise. Testing at only one bitrate setting can never reveal that flattening point, which marks exactly where a product starts paying for data that nobody can actually perceive. In testing across resolutions and bitrates, codecs were found to already be close to their best possible quality by around 2,500 kbps, with H.265 reaching its best 1080p scores at just 500 kbps, meaning additional bitrate beyond that bought almost nothing visible.
Lab numbers are a starting point. Real devices tell you the rest
We measure codec performance and video quality on real devices, real networks, and real battery conditions, not estimates.





