← New search

Other meanings of Fumitada Itakura

Engineering

Fumitada Itakura

Fumitada Itakura (born 1940) is a Japanese electrical engineer and computer scientist whose work on linear predictive coding (LPC) and the Itakura–Saito divergence laid the foundations for modern speech coding, speech recognition, and audio compression. His algorithms enabled efficient digital transmission of speech and became integral to mobile telephony and early speech synthesis systems.

1940
Born
Year of birth
1966
Master's degree
Nagoya University
1975
IEEE Fellow
Elected
1986
Marconi Prize
Co-recipient
1

Early life and education

Fumitada Itakura was born in 1940 in Japan. He studied electrical engineering at Nagoya University, where he received his bachelor's degree in 1963 and his master's degree in 1965. He then joined the Electrical Communication Laboratory of Nippon Telegraph and Telephone (NTT) in 1966, beginning a career that would span several decades. His early work focused on speech analysis and synthesis, driven by the need to transmit speech efficiently over telephone lines. Itakura's doctoral research, completed at Nagoya University in 1975, formalized his contributions to statistical speech modeling, particularly the maximum likelihood approach to linear prediction.1

2

Linear predictive coding and the Itakura–Saito divergence

Itakura's most influential contribution is the development of linear predictive coding (LPC), a method that represents the spectral envelope of a digital speech signal in compressed form. In 1968, he and his colleague Shuzo Saito introduced the Itakura–Saito divergence, a measure of the difference between two spectral densities that became a standard in speech processing. This divergence is derived from the likelihood ratio of autoregressive models and is used in speech recognition, coding, and model evaluation. LPC became the basis for many speech codecs, including those used in early digital telephony and in the LPC-10 vocoder adopted by the U.S. Department of Defense.23

3

Later career and recognition

In 1975, Itakura moved to the United States to work at Bell Laboratories, where he continued his research on speech recognition and coding. He returned to Japan in 1981 and joined Nagoya University as a professor, later becoming a professor emeritus. His work earned numerous honors, including the IEEE Morris N. Liebmann Memorial Award in 1986 and the Marconi Prize in 1986, which he shared with Bishnu Atal for their contributions to speech coding. He was elected a Fellow of the IEEE in 1975 and received the IEEE Jack S. Kilby Signal Processing Medal in 2005. His algorithms remain foundational in modern audio compression standards such as MP3 and AAC.4

4

Lesser-known aspects

Beyond his well-known work, Itakura contributed to the development of the PARCOR (partial autocorrelation) coefficients, which are used in speech coding and are more robust to quantization than direct LPC coefficients. He also proposed the Itakura distance, a simplified version of the divergence that is widely used in dynamic time warping for speech recognition. In the 1970s, he worked on a speech synthesis system that used LPC to generate intelligible speech from text, an early precursor to modern text-to-speech. His collaboration with Shuzo Saito at NTT produced the first practical LPC vocoder, which was demonstrated in 1971 and later influenced the design of the CELP (code-excited linear prediction) codec used in many modern voice communication systems.5

Glossary

Linear predictive coding (LPC)
A method of encoding speech by estimating the spectral envelope using a linear prediction model.
Itakura–Saito divergence
A measure of dissimilarity between two spectral densities, derived from the likelihood ratio of autoregressive models.
PARCOR
Partial autocorrelation coefficients, a representation of LPC parameters that is robust to quantization.
Vocoder
A device or algorithm that analyzes and synthesizes speech, often used for efficient transmission.

Fumitada Itakura's work on LPC and the Itakura–Saito divergence has had a lasting impact on digital communications and remains a cornerstone of speech processing.