Skip to main content
Peter Mc
March 1, 2017

Can digital audio playback be improved by resolutions greater than 16/44.1?

  • March 1, 2017
  • 93 replies
  • 3500 views
Copied from another thread, courtesy of jgatie - full question was:

"Do you believe digital audio, outside of mastering/production techniques, can be improved by playback resolutions greater than 16/44.1?"
    This topic has been closed for further comments. You can use the search bar to find a similar topic, or create a new one by clicking Create Topic at the top of the page.

    93 replies

    Lyricist III
    March 9, 2017
    Apologies in advance if I don't respond to many other queries.

    why have you not requested a retraction?
    Short answer is that I do request corrections when anything I publish contains errors. But I wasn't the author of this, it was just a press release, 'important advantage' sounds to me like opinion rather than a factual claim, I can see how it could be used as a layman's phrasing of 'statistically significant', and scouring the internet would be impossible.


    “I have [yet] to come across even one for hi res that does a decent job of doing level matched blind AB, leave alone a full protocol ABX… may I be pointed to even one ABX done in line with well established principle” – see the paper.

    I have seen the paper. But is there any way to see a fully documented test itself? A link to one perhaps - that is the only way to it for most of us.
    Unless I read the documented test, I can't conclude to how good it is. I may still not be able to do that, but it will be more than anything I have been able to find till now. By good, I mean how well single variable ABX protocols have been complied with.
    Thank you in advance.


    Most of the studies analysed in the paper were published by the Audio Engineering Society. They are all available from http://www.aes.org/e-lib/ , and can be downloaded freely by AES members.
    One non-AES paper that can be downloaded by anyone is at https://www2.ia-engineers.org/iciae/index.php/icisip/icisip2013/paper/viewFile/160/146 . I don't by any means claim that this is the best, and it used an AB style paired comparison rather than ABX, but it should give you an idea of what a lot of the studies were like.
    March 9, 2017
    Thank you! The Japanese test seems to have definitive conclusions and I will read the methodology in detail to understand how it was done.
    jgatie
    March 9, 2017
    Apologies in advance if I don't respond to many other queries.
    Short answer is that I do request corrections when anything I publish contains errors. But I wasn't the author of this, it was just a press release, 'important advantage' sounds to me like opinion rather than a factual claim, I can see how it could be used as a layman's phrasing of 'statistically significant', and scouring the internet would be impossible.


    I'm sorry, but this is stretching the bounds of reality. No matter how many semantic arguments you wish to put forth, not even a layman would equate 'statistically significant' with 'important advantage'. I stand by my point; stating this study showed a 'important advantage' isn't justifiable marketing spin, or a translation to layman's terms; it is a flat out lie.

    And If I were quoted as such when my results said nothing of the sort (in fact they say quite the opposite), I personally would request a retraction. I don't care it it was presented as an opinion, a fact, or an edict from God. I've seen no retraction, so I am left to think that you agree with the "layman's phrasing" which even to a layman, is a gross misrepresentation of the facts. This is pure marketing BS, and you seem to be a willing participant in that BS.

    In short, it is astonishing that you would blame a mere misunderstanding, a marketing speak translation, or a dumbing down for laymen for what is in reality a lie, attributed directly to you. I'm left to wonder if AES's apparent support of all things Hi-Res has any bearing on your agreement with this "layman's phrasing".
    March 10, 2017

    One non-AES paper that can be downloaded by anyone is at https://www2.ia-engineers.org/iciae/index.php/icisip/icisip2013/paper/viewFile/160/146 . I don't by any means claim that this is the best, and it used an AB style paired comparison rather than ABX, but it should give you an idea of what a lot of the studies were like.

    As you say, there are some issues with this report, and someone more experienced with ABX testing will do a better job of identifying these than I can. And what I point out may seem like nit picking to some, I believe it isn't for the following reason:

    Achieving 90% of all that is needed to eliminate Bias does not mean there is only a 10% chance that Bias will colour the outcome - given how powerful and universal human biases are, that possibility can still remain as high as 100%. It really is necessary to come as close to a 100% elimination of bias to have a dependable result.

    To start with the issues I see: first, the purpose of the study is such that tester bias cannot be ruled out.

    Second, the test says it is a blind test, but I don't see that it is a double blind one. It isn't double blind AB, leave alone not being double blind ABX. Not being double blind, tester influence on the results cannot be ruled out.

    I haven't been able to get enough clarity on the outcome that approx 57% prefer HR to CD, although it is clear that an outcome of no preference for either - as in both sounding the same - isn't possible to be obtained by the forced preference design of the test. Do 43% prefer CD to HR? I find that hard to swallow too! And do these numbers delivery any statistically valid outcome?

    My understanding of human memory of sound quality is that it degrades in seconds and becomes unreliable. To overcome that, a rapid - less than a second - changeover between the tested samples is required to get a valid comparison outcome. This hasn't been employed.

    Two important aspects DO seem to have been well addressed - mastering variability, and sound level matching.

    So, while I have pointed out what my limited understanding tells me are flaws in test protocol, a robust peer review would do better in this effort, I admit, but would probably need access to the testers.

    I don't plan to subscribe to AES, but if anyone can pick and describe any test there that is a robust ABX to determine the audible value of HR, it would be very interesting; till then, I will still maintain that I haven't seen any test that adequately rules out bias and takes into account human auditory issues including memory duration, and proves a preference for HR.

    And if as you say a lot of the studies were like this, I am afraid that any meta analysis of these is just as flawed in its conclusions. To use a relevant analogy, no amount of subsequent effort and technology can overcome the defects in the source master.
    Lyricist III
    March 10, 2017

    It really is necessary to come as close to a 100% elimination of bias to have a dependable result... tester bias cannot be ruled out.

    Good point and I fully agree.

    I don't see that it is a double blind one.
    Another hi-res evaluation paper by some of the same authors stated 'an operator was in a blind area outside the vehicle. It was not a double-blind experiment, but the subject and the operator could not communicate each other.'
    I spoke with one of the authors for clarification on other aspects, and they used similar methodologies in both papers. So its likely only single blind, but with significant effort to reduce tester influence.

    I haven't been able to get enough clarity on the outcome that approx 57% prefer HR to CD ah, that's the nature of these sorts of perceptual evaluation tests. They are actually well established throughout psychophysics (and applied to things like taste tests, perfume preference testing etc.), but its not your typical experimental and control groups as in evidence-based medicine. Think of the control as what would happen if participants were guessing. Then the outcome would be 50%. Note that a 50% result in a preference test with lots of trials does not necessarily imply that those participants can't discriminate, since it could be, for instance, that everyone can perceive the difference but half the participants prefer A and half preferred B. But a 100% result in those circumstances would be extremely unlikely to occur by chance.

    For this experiment, standard error of the mean slightly overlapped 50%, so not considered significant. But it had a small number of participants and trials.

    My understanding of human memory of sound quality is that it degrades in seconds and becomes unreliable - yes, and that would weaken any ability to discriminate. On the other hand, there is also the effect of auditory sensory memory. Studies suggest that it actually persists for quite some time. If the higher resolution content is played first, then one might retain the perception of that content if the lower resolution is played after (assuming that there's anything there to be perceived). And since we are looking at the limits of perception, long samples might be needed for listeners to identify differences. I grouped studies into those with short samples / quick changeovers vs long samples / long changeovers and the latter had much stronger results. But I wouldn't read too much into this since other differences between studies might have been the cause of this.


    I haven't seen any test that adequately rules out bias and takes into account human auditory issues including memory duration, and proves a preference for HR. Even if people can hear a difference and even if that difference is due to higher quality, I doubt you'll find rigorous and general proof of preference. Such is the nature of subjective testing at the limits of perception and dealing with real world conditions. I'll give a couple of examples. In one study, participants were asked which audio format was best quality, and the live feed was the one most often ranked worst, suggesting that perhaps people didn't associate quality with purity. And a classic 1956 study suggested that those who have grown up listening to low quality reproductions may develop a preference for it.

    if as you say a lot of the studies were like this, I am afraid that any meta analysis of these is just as flawed in its conclusions.
    This is the big challenge with any meta-analysis, but luckily there is a lot of good advice from meta-analysis experts on how to deal with it. When I first started looking into this, it became clear that potential issues in the studies was a problem. One option would have been to just give up, but then I'd be adding no rigour to a discussion because I felt it wasn't rigourous enough. And its the same as not publishing because you don't get a significant result, only now on a meta scale. So I set some ground rules. I committed to publishing all results, regardless of outcome. I included all possible studies, even if I thought they were problematic (but discuss in detail potential biases). I decided that any choices regarding analysis or transformation of data would be made a priori, regardless of the result of that choice. And I did sensitivity analysis looking at alternative choices (like whether to exclude studies or analyse them differently).
    One interesting thing that I found, which I did not at all expect, was that most of the potential biases would introduce false negatives. That is, most of the issues were things like not using high res stimuli, or having a filter in the playback chain that removed all high frequency content, or using test methodologies that made it difficult for people to answer correctly even when they heard differences.
    March 10, 2017

    And since we are looking at the limits of perception, long samples might be needed for listeners to identify differences. I grouped studies into those with short samples / quick changeovers vs long samples / long changeovers and the latter had much stronger results.

    Such is the nature of subjective testing at the limits of perception and dealing with real world conditions.

    I take your points; a couple of responses to the above quoted:
    1. What about long samples/quick changeovers? You haven't addressed this method - would this not be better than either of the other two you mention?
    2. The limits of perception reached by CD quality sound for it to be heard to be different/worse than HR and real world conditions that further compound this issue are not often considered by HR advocates. And far too often this difficulty is addressed by HR advocates by having a second variable in the mix, the HR master. Because that can often make the difference night and day, and this reason for that to be the case isn't easily picked up by the layman/market.
    Ryan S
    Retired Sonos Staff
    March 10, 2017
    Hi everyone, and @joshr thanks for joining us! This is a great discussion, and I'd just like to remind everyone to respect the rules of the community and withhold from personal attacks or letting the heat of the discussion make you forget that this is a welcoming place where everyone is entitled to their opinion. Let's keep it friendly!

    Thanks!
    Mark good posts by pressing the like button, and select the best answer on questions you've asked to help others find solutions.
    Lyricist III
    March 11, 2017
    Hi Kumar,
    1. Would be quite interesting. I think the optimal duration of timing and intervals would depend on which cognitive processes play the strongest role for a given task. But anyway, I don't think there was sufficient data for me to do a proper analysis on this. A lot of studies simply didn't give a lot of information on these durations.
    And training seemed to be such a strong variable that it appeared to override everything else. Essentially (and I'm being loose with the language here), every study where participants were carefully trained in the task showed a strong ability to discriminate, every study where participants were untrained did not. And
    Lyricist III
    March 11, 2017
    ... which made the effect of other factors like duration seem weak at best.

    As for 2, you would clearly be comparing something very different if you look at commercial HD and CD quality recordings, if anything else was done in mastering besides just sample rate and bit depth conversion. But this argument works both ways of course. If the HD version is just an upsample of CD quality content, then its highly unlikely that there would be significant perceivable differences. Luckily, almost all studies avoided anything like this, and potential biases were noted and discussed when any such issues could have occurred.
    March 11, 2017


    And training seemed to be such a strong variable that it appeared to override everything else.


    If training is guidance on a range of parameters to be observed/heard, with words suggested for the two ends of the scales for each, it seems to me that it would be useful; anything more starts running into contested territory!
    But to my thinking, the forced preference is the bigger issue - and I would propose that it shows up in the CD V HR results where roughly there are an equal amount that prefer one over the other. More than anything else, I suggest that this shows there is little to choose between the two and people say - if I have to pick one, I pick this; and the fact that preferences expressed in that manner fall equally on either side of the divide has a message to convey about the nature of the outcome.