What does 3D Audio actually mean?

Looking at Ambisonics from a different angle.

Text: Elettra Bargiacchi Interview Partner: Johann-Markus Batke Pictures: Johann-Markus Batke

Block 1

In the previous issue (VDT-Magazin 2026/3) we discussed Wave Field Synthesis, and today we turn to Ambisonics – the other face of the same research effort: reconstructing a believable, three-dimensional sound field.

My partner in this conversation is Johann-­Markus Batke, Professor at the University of Applied Sciences Emden/Leer and formerly Senior Scientist at Technicolor, where his research centered on HOA and binaural synthesis. He’s also a friend and colleague within the VDT R&D Department, which makes it a pleasure to dedicate this article to our exchange.

Block 2

Johann-Markus Batke is Professor of Communications Engineering at Emden/Leer University of Applied Sciences, where his work centres on digital signal processing. He lectures across the Media Technology, Electrical Engineering, Media Informatics, and Industrial Informatics degree programmes. His primary research interests lie in 3D audio and musical acoustics. In addition to his academic duties, he serves as a member of the management team for the VDT R&D Department.

Block 3

We started with a simple question. What does 3D audio actually mean? And why isn’t stereo enough? What is 3D Audio at all? This question, philosophical in a way, has accompanied Batke’s research from the very beginning – the technical side of Ambisonics only makes sense once you’ve sat with it.

Nobody was really asking this question, people were and still are used to stereophonic representation of audio content. But stereo is just an aesthetic image of what you’d like to show in the spatial sense, it’s more or less a kind of audio art. You are dealing with spatial audio information, but you do not represent it in a way that conforms to the physical audio scene. 3D audio, on the other hand, represents audio in a more physical manner, depicting the real acoustic field as it is.

What is 3D Audio at all?

Seen this way, Ambisonics and WFS are two technical answers to the same underlying question: what do you really need to recreate a sound field? WFS answers by physically recreating the wave field itself in a 2-dimensional plane. HOA (Higher Order Ambisonics) instead describes the sound field entirely in three real dimensions using spherical harmonics, with a spatial resolution that depends on the order: the higher the order, the more spatial details are delivered, and transmitted coefficients are required.

You can imagine that process as a blurry image. If you take just first-order Ambisonics, you have a very blurry image of the entire soundscape, if you want more spatial detail, you increase the order. The higher the order, the more signals and the less blur.

Block 5

Johann-Markus Batke‘s Lab with 22.2 system
Johann-Markus Batke‘s Lab with 22.2 system

So, put simply, what’s the main difference between Wave Field Synthesis (WFS), Ambisonics, or other 3D audio methods?

WFS is basically limited to the description of a two-dimensional plane, while HOA gives a different mathematical description of the entire sound field in all three dimensions, meaning also height. Another difference is that with HOA we also have a consistent recording approach: if we use a Higher Order Ambisonics microphone, we can capture the entire scene, recode it in the HOA format and have a three-dimensional representation immediately.

The problem is then to tweak it for production – that’s where dedicated plugins, like the one developed by the IEM Graz, come in because you need tools to shape the soundscape in a very specific way. With all other approaches, you have to record audio objects, know their position and also know the room acoustics – Ambisonics, instead, offers an immediate, self-contained recording format for the whole 3D scene.

Ambisonics and Binaural: A Perfect Match

This issue of the Magazin is dedicated to binaural audio, a topic that pairs naturally with Ambisonics: Johann-Markus sees Ambisonics as an intermediate representation from which binaural rendering can be derived.

If you use head tracking and have no preferred listening direction, you need all listening directions at once – that’s what makes Ambisonics very feasable. With binaural synthesis, you’re relating all directions to your two ears, and if you move your head, that relation changes.

Head tracking removes any fixed listening direction, so you need a format representing every direction simultaneously – exactly what Ambisonics provides, ready to be rendered binaurally as the listener’s head moves.

What’s Next for Ambisonics?

Looking ahead, you don’t expect ­Ambisonics itself to change dramatically?

I would expect not the Ambisonics itself, but its production tools and user interface to probably develop further.

Conclusion

Ambisonics and Wave Field Synthesis are two different paths toward the same goal: a physically faithful reconstruction of 3D sound. But underneath both, beyond the technicalities, the real question Batke keeps revisiting is profound: what is the true meaning of representing an actual physical sound field rather than just creating an aesthetic image of it.

Ambisonics offers a compact, format-agnostic description of the whole sound scene, flexible enough to be recorded, shaped in production, played back on different setups, and rendered in binaural for headphones. As head tracking and immersive formats spread, it’s this flexibility, more than the underlying mathematics, that seems destined to keep evolving – through better production tools and simpler interfaces for the sound designers of tomorrow.

Block 9

Elettra Bargiacchi is an Italian sound designer, composer, and musician based in Leipzig. Trained in Classical Guitar and Composition (Conservatory of Milan) and Audio Post Production (Abbey Road Institute, London), she is passionate about immersive audio and innovation, works on international film and podcast productions and research, has authored papers on next-generation audio, and is a member of the VDT R&D Board.