Ph.D. program:

  • “Didactic Application of Extended Reality for Drumming and Polyrhythms,” J. Pinkl (2025).
    This dissertation is primarily devoted to both furthering research in the field of XR- based motor learning tools and to make polyrhythms more comprehensible for musicians and nonmusicians. The contributions of this work include: “VR Polyrhythmic Samchillian,” this is a VR-based musical instrument with controls based on the Samchillian, an instrument invented by Leon Gruenbaum in 1988. The Samchillian is unique in its delta-based control, which gives users access only to buttons that modulate the pitch of the tone last played in set intervals, rather than buttons that correspond to absolute pitches. The VR-based instrument is equipped with a similar interface allowing users to build upon and modify polyrhythms within the virtual scene. As its use is unbounded by musical skill of the user and not based on rhythmically limited grid-based programming of standard drum machines, the VR Polyrhythmic Samchillian allows accessible interaction with a wide array of two-voice polyrhythms. VR tool to teach drumming using Action Observation- and Virtual Co-embodiment-based learning. This tool allows users to occupy a virtual scene with an autonomous agent exemplarily demoing rudimental and polyrhythmic drumming exercises. Users mimic the movements of the exemplar in the Action Observation-based scenes, and share the performance of the rhythm with the exemplar in the Virtual Co-embodiment scenes. In the latter, the author coined this single-user approach to Virtual Co-embodiment as “halvatar.” Although existing within Virtual Co-embodiment’s definition and scope, to the author’s knowledge, the realization of this concept for didactic motor learning did not heretofore exist. Multimodal MR tool to teach drumming using Action Observation-based learning. This contribution is an extension of the aforementioned VR tool, and its most noteworthy advancements are haptic functionality and video see-through. Users receive vibrational reinforcement of rhythmic exercises and the physical electronic drums pads, in addition to their virtual counterparts, remain visible in a mixed reality scene with overlaid elements. Data on rhythmic accuracy improvement of subjects using Multimodal MR-based Action Observation learning vs. video-based learning. Past studies on VR-based Action Observation have shown mixed results on such tools’ effectiveness for learning. For example a study with a VR-based Action Observation tool to teach prosthetic limb control showed significant improvement when compared with a control group, whereas a study with a comparable tool to teach tai chi had the opposite observed effect. In addition, Action Observation, VR-based and otherwise, in musical applications lacks data from quantitative studies in general. After finishing development of the MR tool mentioned above, a 20-subject study was conducted to compare improvement of timing accuracy of rhythmic performance between two groups. An experimental group, using the author’s system to practice, and a control group, practicing drumming via video demonstrations, had the improvement of their timing errors compared and results showed a significantly greater decrease in error in the experimental group—239 ms (z-ratio= 3.520, p < 0.001). These results demonstrate potential of such mixed reality tools in musical applications while also contributing evidence that MR-based Action Observation can be an effective practice method for drumming.

    Source code


  • “Enhancement of Audio Directionalization in Virtual Environments,” C. Arevalo (2024).
    This study focused on enhancing the efficiency and subjective performance of audio spatializers in extended reality (XR) applications. Our approach centers around the compression and personalization of Head-Related Transfer Functions (HRTFs) to address limitations in current audio spatializers. We developed an audio spatializer capable of rendering all audible audio sources within a current game engine by using over one million HRTFs compressed via Eigen decomposition. While our proposed method showcases improvements in processing time and memory requirements compared to existing methods, challenges remain in accurately reproducing auditory locations as judged by human subjects. This discrepancy is likely attributable to mismatches between generic HRTFs, commonly used in audio spatializers, and those of a particular listener. To improve the subjective performance of audio spatializers, we propose a method for the personalization of HRTFs through the manipulation of latent space features extracted from multiple HRTF databases using autoencoders. These are a type of Artificial Neural Network that efficiently reduces data dimensionality. The research also encloses the development of a tool to collect subjective evaluations and assess the subjective accuracy of the implemented solution. Subjective experiments show that the personalization method has better performance for elevations located at the back compared to a non-personalized HRTF database. By addressing the compression and personalization of HRTFs, this study aims to enhance spatial audio experiences in XR environments.

    Source code


  • “Applications of Evolutionary Algorithms to Digital Audio Signal Processing,” E. Ly (2024).
    With Artificial Neural Networks (ANNs) being at the forefront of Artificial Intelligence (AI) research in recent decades, other Machine Learning (ML) models within the larger field of AI can sometimes be overlooked as viable alternatives for solving complex problems in many fields. This dissertation examines some possible use cases of Evolutionary Algorithms (EAs) by applying them to audio Digital Signal Processing (DSP) problems in particular, including those that have previously seen many proposed solutions using ML. We start with the creation of Room Impulse Responses (RIRs) that can be used within a real-time convolution reverb audio effect plugin. Using an Evolutionary Programming (EP) approach, a user is able to gain some control over the shape of the resulting RIRs through various parameters defined in the ISO 3382-1 standard (e.g., reverberation time, early decay time, and clarity), the values of which determine the fitness of potential RIRs. The resulting rooms can range from those whose RIRs can be recorded in the real world, to virtual spaces that may be physically impossible to represent. Results from a subjective evaluation show that such perceptual differences were reduced when the EA was executed for a sufficient number of generations, or when the input audio signals consisted of only speech. We then proceed to a more general case where entire software synthesizers are generated via Recurrent Cartesian Genetic Programming (RCGP): given a target audio signal, DSP programs expressed as directed cyclic graph structures are evolved to generate an approximation of the target audio with arbitrary accuracy. The candidate program with the smallest error, using a fitness function based on Mel-Frequency Cepstral Coefficients (MFCCs) quantifying perceptual differences, is returned as the best fit solution. After an experiment evaluating the effects of several RCGP parameters, we determined that a classical Cartesian Genetic Programming (CGP) model with a weighted function set that accounts for prior knowledge was able to generate steady state signals with minimal error. This method was eventually generalized even further to include the generation of DSP audio effect chains, with the Log-Spectral Distance (LSD) being used as a fitness metric to compare the frequency spectra of their respective impulse responses. We then evaluated our method by generating various Infinite Impulse Response (IIR) filters, with the accuracy of these filters being dependent on the order of the target filters to be replicated. Working prototype implementations for these methods are publicly available as free software for demonstration purposes.

    Source code

    Demonstration


M.Sc. program:

  • “Head-Related Transfer Function Compression Using Splines,” T. Krüger (2026).
    This work presents a spline-based method for compressing and reconstructing Head-Related Transfer Functions (HRTFs) that preserves perceptual quality. The method operates on the magnitude spectrum of minimum-phase head-related impulse responses (HRIRs) through a four-stage pipeline: (1) acquiring the minimum-phase HRIRs, (2) transforming them into the frequency domain and applying adaptive Wiener filtering, (3) extracting a minimal set of control points through derivative-based analysis, and (4) reconstructing the spectrum with piecewise cubic Hermite interpolation (PCHIP). Across 301 subjects from the SONICOM database, this pipeline reaches a compression ratio of 2.8:1 with spectral distortion ≤ 1.0 dB in every ERB and a mean absolute Interaural Level Difference (ILD) error of 0.10 dB. The strict 1.0 dB threshold yields a variable, measurement-specific number of control points that cannot be compared across subjects. To obtain a standardized, editable representation, we relax this threshold and fix the budget to exactly 43 control points per channel, keeping the single most prominent spectral feature in each of 42 ERBs plus one anchor point. Reconstruction from these 43 points is performed with Akima interpolation, which only needs the stored points, and with a hybrid method that selects the best spline type (PCHIP, Akima, or thin-plate spline) per segment from the original spectrum, serving as an upper bound. Across 302 subjects, this fixed representation reaches approximately 3.0:1 compression at a mean spectral distortion of 3.94 dB (Akima) and 3.55 dB (hybrid). Whether this increased objective distortion is perceptible was tested in a localization experiment with 18 participants, using one of the worst-reconstructed subjects in the dataset. No significant difference in localization accuracy was found between the original and reconstructed HRTFs for either azimuth (p = .311) or elevation (p = .847). The per-location analysis showed the same error pattern across all 16 source locations regardless of method. The Akima reconstruction matched the uncompressed original despite its higher objective distortion, while the hybrid offered no perceptual advantage. These results indicate that preserving the most prominent spectral feature per ERB is sufficient to maintain the spatial cues needed for localization, and that the common 1.0 dB distortion threshold could be more conservative than localization requires. Lastly, the fixed 43-point format provides a standardized, editable foundation for future HRTF personalization.

    Source code

  • “Prediction of Non-Modal Phonation From Auditory Features,” Y. Sakai (2026).
    We investigate automatic phonation prediction using machine learning classifiers, using features derived from auditory models. Previous studies have relied on mel-frequency cepstral coefficients (MFCCs) or acoustic parameters such as cepstral peak prominence, harmonics-to-noise ratio, and spectral tilt. Auditory models simulate the transformations the acoustic wave undergoes in the human auditory periphery. Thus, their outputs provide representations that align more closely with human perception of phonation than conventional acoustic measures. We evaluated six neural network architectures: BiLSTM, CNN-BiGRU, CNN-BiLSTM, Transformer, CNN-Transformer, and 2D CNN. These classifiers were trained on MFCCs and compared against three auditory-model feature sets: Gammatone cepstral coefficients (GTCCs), the neural activity patterns (NAPs) output of the Cascade of Asymmetric Resonators with Fast-Acting Compression (CARFAC) cochlear model, and the adaptation stage output of the Computational Auditory Signal Processing and Perception (CASP) model. In addition, we compared these architecture–feature combinations against classifiers trained with pre-trained SSL models (HuBERT, Wav2Vec2, and WavLM), which were fine-tuned directly on recorded speech. We conducted both multi-class classification across five phonation types and binary (one-vs-rest) classification for breathy, creaky, and modal voice, using a dataset containing eight languages. For binary classification, GTCC features achieved the highest macro F1 scores of 73.79% (breathy, CNN-BiLSTM), 67.37% (creaky, BiLSTM), and 81.34% (modal, Transformer). CASP auditory features approached GTCC on breathy (70.15% F1, CNN-Transformer) and modal (78.07% F1, CNN-Transformer). For multi-class prediction, the best F1 was achieved by GTCC (57.78%, Transformer), followed by CASP (56.06%, CNN-Transformer), MFCC (55.55%, CNN-Transformer), and CARFAC (53.07%, CNN-Transformer). CNN-Transformer was the only architecture to rank among the top two across all four features. While WavLM Large (64.04% F1) and HuBERT Large (63.59% F1) surpassed feature-based architectures in multi-class classification, Wav2Vec2 Large and all base-size models collapsed to near-chance performance (F1 below 10%). As a case study, we applied this framework to the detection of Parkinsonian speech. A CNN-BiGRU model with MFCCs achieved the best classification accuracy at 72.06%, while the same architecture trained on CARFAC NAPs reached 71.17%, approaching the MFCC baseline. These results suggest that biologically inspired representations capture perceptually relevant phonation information and are viable alternatives to conventional cepstral features, though their relative performance depends on the specific architecture and task; however, biological fidelity alone did not guarantee superior performance.

    Source code (feature extraction)
    Source code (phonation prediction)

  • “Data Musicalization Using AI-Generated Music,” T. Odaira (2026).
    We propose an emotion-driven sonification method that employs AI-generated music to convey temporal data. Conventional sonification techniques typically rely on direct mappings from data values to acoustic parameters such as pitch, timbre, or tempo, which often require prior training for effective interpretation. To address this limitation, the proposed method introduces valence and arousal as an intermediate affective representation between numerical data and sound. In the proposed system, normalized continuous data are mapped onto a two-dimen\-sional valence--arousal space, from which an emotional state is derived and used to condition music generation via the Suno framework. A complete data-to-music pipeline was designed and implemented, including data normalization, affective mapping, melo\-dy generation, and music composition, and was made accessible through an interactive interface. A subjective evaluation was conducted using real-world radiation data to assess whether the generated music could convey data-dependent information to untrained listeners. The results showed a modest but measurable rank correlation between participants’ perceived happiness rankings and ground-truth rankings, despite no explicit definition of happiness being provided. This finding suggests that affective cues embedded in AI-generated music can support the communication of data trends without requiring listener training. While the evaluation focused on a single emotional dimension, the results provide initial evidence for the feasibility of emotion-driven sonification. Future work will extend the evaluation to multiple affective dimensions and investigate more systematic and automated mapping strategies between data structures and the valence--arousal space.

    Source code


  • “Towards automatic declipping of audio recordings of speech using deep learning techniques,” Wen Wen (2025) .
    Clipping is a form of audio distortion that occurs when a signal's amplitude exceeds the dynamic range of a system, leading to loss of information. Declipping is a process that seeks to recover the original signal with the available information. Existing methods such as ASS-PEW, while effective at predicting clipped samples, are computationally expensive preventing their usability in real-time applications. In addition, this method does not consider the amplitude of the original signal and instead relies on other processes to scale the output. This research explores using Long Short-Term Memory (LSTM) neural networks to develop a declipper specifically for spoken words, with the goal of minimum processing delay. Compared to existing methods, our resulting pre-trained model shows promise in restoring the amplitude of the signals but falls short of reconstructing the clipped samples, in its current state it could be used as a complement to existing methods. In future research, we will consider creating sub-models that specialize in reconstructing the clipped samples to improve our current system.

  • “Prediction of psychoacoustic roughness using machine learning,” S. Yoshida (2024).
    Psychoacoustic roughness is a subjective value that expresses the sense of roughness of sound and is one of the sound quality evaluation indices along with loudness, sharpness, and fluctuation strength. Generally, a sound with high roughness value induces discomfort in the human auditory sense. The purpose of this research is to develop a model to predict roughness using machine learning. Although several existing prediction models have been proposed, development of a prediction model with higher accuracy and wider applicability is expected to be useful in sound quality evaluation and acoustic design. Two methods, CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network) with Bidirectional LSTM (Bidirectional Long Short Term Memory), are used as machine learning methods. Out of a total of 54135 pieces of data, 80% were used for training and 10% each for testing and validation. The data used to train the model are values obtained from subjective experiments in previous research. After training, we compared the predictions of the machine learning model with those of existing standardized models for roughness prediction and actual subjective values to examine the accuracy and usefulness of the machine learning model and future issues. Models trained with CNN yielded more accurate results than standardized models in most cases, while RNN were less accurate than CNN and in some cases less accurate than standardized models.

    Source code


  • “Development of ad hoc loudspeaker arrays based on smart devices,” K. Fujita (2024).
    We present an ad hoc loudspeaker array, loudspeakers and sound sources move arbitrarily on the horizontal plane in real-time. The spatializer, derived from the "room within a room" spatialization method, comprises a listening area defined by the convex hull formed by the loudspeakers and a larger area (defined by a user) where audio sources can move. The ad hoc loudspeaker array was developed as an iOS application. To precisely determine the locations of the loudspeakers, the application used an iOS library designed for indoor positioning technology that employs Ultra-Wide Band. Peer-to-Peer communication facilitates exchange of sound and loudspeaker location information. Subjective evaluation was conducted to assess the application's capabilities, featuring five distinct patterns of sound and loudspeaker movements. The task of a participant was to listen to sounds and to rates on a 5-point Likert scale, how much they agreed with given statements. The results revealed consistently positive ratings for moving stimuli, irrespective of loudspeaker movement. However, the results for stationary stimuli were less satisfactory. These results suggest that spatialization performance can be improved by increasing the number of adapted loudspeakers.

    Source code


  • “Effect of local spectral tilt modifications in the intelligibility of speech in noise,” Y. Sato (2024).
    We compared spectral tilt modifications of plain speech made by Linear Predictive Coding—LPC transplantation and fractional roll-off filtering—FRF. These modifications were done at utterance-, phone-, or frame-level to equate the spectral tilt of the same utterances produced in noise (Lombard speech) by the same speakers. When mixed with speech-shaped noise at the same Signal-to-Noise Ratio (SNR), speech treated with LPC transplantation yielded larger objective intelligibility gains relative to those of speech treated with FRF. However, it also yielded larger spectral tilt errors for frame-based modifications. Subjective evaluation for utterance- and phone-level modification also shows larger correct keyword rate gains on speech treated with LPC transplantation compared with speech treated with FRF. This finding suggests that spectral tilt benefits assigned to Lombard speech may be smaller than previously thought, especially for those reports based on phone-based LPC transplantation.

    Source code


  • “Sonification of patterns in Big Data by machine learning generated music,” K. Diaz-España (2023).
    We propose a sonification method for time-series. Conventional sonifications directly assign different aspects of data to auditory features and the results are not always ``musical’’ as they do not adhere to a recognizable structure, genre, style, etc. Our system tackles this problem by learning orthogonal features in the latent space of a given musical corpus and using those features to create derivative compositions. To train our model, we use an array representation of Musical Instrument Digital Interface (MIDI) files as input. We propose using a Singular Autoencoder (SAE) algorithm that identifies the most important Principal Components (PCs) in the latent space. As a proof-of-concept, we created sonification of Fine Particulate Matter (PM2.5) from the AEROS database. Also, we created sonifications of ionizing radiation measurements obtained from the Safecast project. Although the system successfully generates new compositions by manipulating the latent space, with each principal component changing different musical aspects, these changes may not be readily noticeable by listeners, despite the PCs being mathematically decorrelated. This finding suggests that higher-level features (such as associated emotion, etc.) may be needed for better results

    Source code


  • “Real-time stereophonic sound using a ring of loudspeakers,” S. Fujisawa (2023).
    We present the development and evaluation of a real-time spatialization system based on inverse filtering and loudspeaker grouping. This spatial processing system uses a ring of loudspeakers, grouping the lateral loudspeakers for vertical spatializations. By this grouping, the reproduced sound mimics real elevated sounds which can reach the front and the back of a listener's pinnae simultaneously. This real-time implementation requires low CPU-processing and memory, as demonstrated by objective evaluations. This is partly due to the use of a database comprising more than one million compressed Head-Related Transfer Functions (HRTFs)---most of them interpolated from measured ones, and filtering in the frequency domain. Besides grouping, inverse filtering decreases the influence of each loudspeaker's location on the spatialized sound. We found that this filtering can be replaced by intelligent equalization systems currently used in some high-end loudspeakers (e.g., Genelec 8320A). The proposed system can be used with any number of loudspeakers; we tested the real-time implementation using a quadraphonic layout. In addition to objective evaluations, subjective differences in accuracy when using external equalization instead of inverse filtering were also investigated, suggesting that both filters are comparable. Finally, results of these evaluations are discussed and conclusions are drawn.

    Source code


  • “Data Sonification using Music Genre Latent Space,” Y. Kariyado (2023).
    In this thesis, musical instruments associated with big data are introduced. First, I have implemented three-dimensional cellular automata with an auralization tool. This program makes some unique sequential sounds according to the location of cells and their states. I obtained the concept from this research that it might be possible to automatically generate music by using big data and neural networks. For that reason, I implemented a data sonification tool based on autoencoders. The network consists of two types of autoencoders, Musical Instrument Digital Interface (MIDI) files are used as input data. In addition to MIDI files, Hymn MIDI files are adopted because of the large number of files on the internet, as they have a length and rhythm suitable for learning its structure. Since this program outputs MIDI messages, users can explore big data by using Virtual Studio Technology (VST) and MIDI synthesizers. Associating big data with latent variables can be considered sonification. The purpose of this research is two-fold: Representing data with sonified sound helps all people, not just the visually impaired. In the entertainment scene, a musician who is not familiar with science can interact or communicate data as a sound. In the current research, the implementation of generative music by using a neural network has been done. A network architecture in previous research usually used a combination of recurrent neural network Recurrent Neural Network (RNN) and long short-term memory Long Short-Term Memory (LSTM). That model is a network model for dealing with time series data, and it’s effective for generating music. In conclusion, I obtained that the program could generate sonified data using music genre latent space. However, latent spaces are not entirely intuitive due to the black box of machine learning models. Therefore, the possibility of low reproducibility when using other data was indicated.

    Source code


  • “Offline automatic speech recognition as second language pronunciation training tool,” L. Tang (2021).
    The aim of this research is to help learning a second language by using an offline and mobile Speech Recognition System (SRS). By prompting users to utter a word from a minimal pair, the system can help to identify the most likely word heard by a native speaker, so that users can focus on improving their pronunciation until the uttered words are correctly recognized. We used PocketSphinx as the underlying technology for this project. With this tool, we implemented an offline speech recognition system. The offline part is important since it allows users to practice anywhere, anytime, for as long as they want, without spending mobile network communication data packages which are a valuable commodity in some parts of the world. Our proof-of-concept application is targeted for Android devices. In order to verify whether this system has a positive impact on second language acquisition, Japanese speakers learning English were asked to use our mobile SRS (dubbed CAPTANG) for pronunciation learning. We collected their utterances before using the system and after three weeks of training, and compared the SRS accuracy achieved for each participant. Preliminary results indicate that our approach has a beneficial impact on L2 learning.

    Source code and Android App


  • ” Improving speech intelligibility by harmonic structure modifications,” S. Hirata (2021).
    There are many situations where it is difficult to understand speech because of noise masking effects. For example, inside a running train in a busy city. The goal of this research is finding ways to increase speech intelligibility in noisy conditions without raising the sound level. Specifically, we do that by changing the harmonic structure of an audio signal. We explored which harmonic changes are the most effective to improve speech intelligibility under noisy conditions. In this study, we combined two approaches: harmonic distortion and dynamic range compression. Objective evaluation results showed that full-wave rectified speech had higher intelligibility than half-wave rectified speech and speech processed using Spectral Shaping with Dynamic Range Compression (SSDRC). SSDRC is an algorithm for improving speech intelligibility in noisy conditions that Zorila et al. developed jointly in 2012. It was also confirmed that the intelligibility was further increased by using dynamic range compression. However, in a subjective evaluation, speech processed with full-wave rectification and dynamic range compression achieved lower intelligibility than speech processed using SSDRC.

  • “Accurate spatialization of several near-field sound sources in VR,” C. Arévalo (2021).
    A method to accurately spatialize several audio sources is presented. By using Eigen decomposition, the proposed method reduces the memory footprint of a convolution-based spatializer from 36.30 MB to 1.8 MB (about a 20:1 compression ratio). Synthetic HRTFs in the compressed database were set to have less than 1 dB spectral distortion between 0.1 and 16 kHz. The differences between the compressed measurements with those in the original database do not seem to translate into a degradation of perceptual location accuracy. The high degree of compression obtained with this method allows the inclusion of interpolated HRTFs in databases for easing the real time audio spatialization in VR. Results of objective evaluations indicate that real-time behavior is comparable with spatializers currently used in VR environments.

    Source code


  • “Genetic Reverb: Synthesizing Artificial Reverberant Fields Via Genetic Algorithms,” E. Ly (2020).
    Genetic Reverb, a user-friendly VST 2 audio effect plugin that performs convolution with an audio signal and a synthetic Room Impulse Response (RIR) generated via a Genetic Algorithm (GA), was developed. The plugin can be utilized by audio engineers, game developers, and music producers to create custom RIRs without the need of recording them in real-world en- closures. The parameters of the plugin include some of the room acoustics parameters defined under the ISO 3382-1 standard (reverberation time, early decay time, clarity), among others. With these parameters, the user is given some control over the shape of the resulting RIRs as the fitness values of potential RIRs can be determined from these parameters. In the GA, these RIRs are initially generated via a modified Gaussian noise method, and then evolved via truncation selection, random weighted average crossover, and mutation via Gaussian multiplication. These operations are repeated until a certain number of generations has passed or the fitness value reaches a threshold. Either way, the best-fit RIR is returned. A Binaural Room Impulse Response (BRIR) can also be generated for a binaural reverberation effect, assigning two different (mono) RIRs to the left and right stereo channels. With Genetic Reverb, new RIRs that represent virtual rooms, some of which may even be impossible to replicate in the physical world, can be generated and stored. Through subjective evaluation, it was determined that RIRs generated by the GA were perceptually distinguishable from similar real-world, recorded RIRs, but the perceptual differences were reduced when RIRs obtained from GAs with longer execution times were used, or audio signals consisting of only reverberated speech (as opposed to other sounds with broader frequency ranges) were tested.

    Source code


  • “Study of floor reflection on sound elevation perception in virtual reality,” N. Fukasawa (2020).
    Humans hear both direct and sound reflected on the floor in most common settings. Reflections effects on auditory location accuracy have been shown in previous research to improve the localization of sound sources in the horizon. But, the effect of floor reflections (arguably the most common ones) on elevation perception are still not well understood. In this project, we recorded a set of Head-Related Transfer Functions (HRTFs) with and without floor reflections and compare the elevation subjective accuracy obtained for audio processed with these HRTFs.

    HRIR database in SOFA format


  • “Elevation of sound by spectral energy equalization and delay adjustments using single-layer loudspeaker arrays,” T. Nagasaka (2015).
    We investigated the relative influence of spectral cues on elevation localization of virtual sources. Comparing five methods and two of them based on Vector-Based Amplitude Panning: 3D Vector-Based Amplitude Panning (3D-VBAP), 2D-VBAP in conjunction with HRIR convolution. Three of them are equalizing filters method which filtered the stimuli to simulate spectral peaks and troughs naturally occurring at different angles, the modification of equalizing filters (delay adjustments), and real positions. A single horizontal loudspeaker array was used for three methods (2D-VBAP in conjunction with HRIR convolution, and equalizing filters with/out delay adjustments). The experiment was divided into two experiments. In the former experiment, smallest absolute errors were observed for the 3D-VBAP judgements regardless of azimuth; no significant difference in the mean absolute error was found between the other two methods in a first experiment. However, for most presentation azimuths, the equalizing filter method yielded the least dispersed results. In latter experiment, an improvement of localizations was observed in equalizing filters treatment by a change of HRTF database and an addition of delay adjustments. We found the changes were related to reproduction of elevated sounds, however the relationship between localization and these changes such as adjustment was unclear. These results could be used for improving elevation localization in two-dimensional VBAP reproduction systems. In the whole of the experiments, we developed an experiment system that use Pd as an iOS application, and OSC protocol was used as the back-end system. Most of implementation process were GUI based programing, so it was understandable for an experimenter who was not familiar with character base programing. The system had a possibility to reduce time cost that was caused by previous methods such oral communication and a notes taking.

  • “Lateralization of sound by spectral energy equalization and delay adjustments using single-layer loudspeaker arrays,” S. Nogami (2015).
    In this research, we investigated influence of spectral energy changes and delay on azimuth judgements. First, we investigated influence of spectral energy changes. So, we compared subjective judgements of azimuth obtained by three methods: Vector-Based Amplitude Panning (VBAP), VBAP mixed with binaural rendition over loudspeakers (VBAP+HRTF), and a newly proposed method based on equalizing spectral energy. In our results, significantly smaller errors were found for the stimuli treated with VBAP+HRTF; differences between the other two treatments were not significant. Regarding spherical dispersion of the judgements, VBAP results have the greatest dispersion, whereas the dispersion on the results of the other two methods were significantly smaller, however similar between them. These results suggest that horizontal localization using VBAP methods can be improved by applying a frequency dependent panning factor a opposed to a constant scalar as commonly used. And we hypothesized that including Interaural Time Difference (ITD) is efficient on azimuth judgements. Secondly, we investigated influence of delay. So, we compare subjective judgements of azimuth obtained by three methods: reproducing the stimuli from the real loudspeakers, ambisonics, and a newly proposed method based on equalizing spectral energy including delay. In our results, adding delay has a beneficial effect on accuracy of panning judgements. Our method yielded smaller absolute errors than ambisonics. Moreover it yielded smaller absolute errors than our previous method. These results suggest that including delay adjustments can improve horizontal localization.

  • “Relative influence of spectral bands in horizontal-front localization of white noise,” T. Sugasawa (2014).
    In this research, important frequency ranges to recognize sound as coming from the front are investigated and a useful method to reproduce realistic sound for that direction in the absence of front loudspeaker is suggested. Stereophonic systems are reproducing sound systems that use two-channels, usually arranged symmetrically on middle-plane and at the same distance from listener’s position. With such a system is possible to present realistic sound in frontal direction, but its localizing accuracy and sound spreading are worse than multichannel surround system since there is no front loudspeaker. In order to ameliorate this, I focused on manipulating the frequency spectrum and investigated the relationship between energy in spectral bands and front localization. Yunoue [1], who graduated from the University of Aizu, had investigated same contents previously using three loudspeakers. He divided frequency range between 0.02--22.05 kHz into 13 bands and conducted experiments to determine which bands are more important to localize sound images in the frontal area. However, there are some problems in his method that may affect his results. To assess the correctness of his findings some modification were introduced and the same experiment was conducted again. In addition to that experiment, data from Head-Related Transfer Functions (HRTFs) was analyzed to create compensation filters on 0° and ± 30° to improve focus of front sources in stereophonic systems. The created filters were convolved with some sound sources and reproduced via two-way loudspeakers. In addition, as a way of comparison, sounds applied panning technique were also reproduced. The results show that the method of using inverse filters was better than simple panning at improving the perceptual focus of the frontal image. This method could be easily implemented in real time systems and probably extended to other spatial dimensions.

B.Sc. program:

  • "Real-time Overtone Singing Audio Effect," M. Fujiwara (2026).
    Overtone singing is a vocal technique that produces salient harmonics in conjunction with the fundamental frequency, resulting in the perception of two tones. The mastery of this technique requires prolonged training. Therefore, the objective of this study is to develop a speech processing system that reproduces this complex singing technique for expressive purposes. An analysis was conducted to assess the frequency, filter quality, and intensity of each formant in overtone sung sounds. The findings from this analysis were then employed in the system design. The system developed utilizes normal singing voice as input, subsequently reproducing overtone singing through the selective emphasis and control of overtone components within the range of +24 to +35 semitones by employing band-pass (bp) filters. Overtones are selected in real-time by users employing a virtual musical keyboard. A demonstration of the implemented system is offered as a way of validation.

    Source code


  • "Breathy Speech Recognition Using Auditory Models and Machine Learning," I. Saito (2026).
    To assess the advantages of auditory models in classifying breathy phonation, this research compared the performance of models trained with Gammatone Frequency Cepstral Coefficients (GFCCs) against those trained with conventional Mel-Frequency Cepstral Coefficients (MFCCs). We utilized a dataset comprising five languages with phonemic phonation contrasts and a Long Short-Term Memory (LSTM) network with an attention mechanism was implemented for binary classification. Evaluation using 5-fold cross-validation demonstrated that the GFCC-based model achieved a mean balanced accuracy of 81.83%, outperforming the MFCC-based model (79.77%). The GFCC-based model reached a recall of 70.95%, an improvement of approximately 3.6 percentage points over the baseline. These findings suggest that auditory features are effective for capturing the subtle acoustic features of breathy speech. The results also highlighted the challenges posed by class imbalance, indicating that further refinements are needed to enhance detection sensitivity across diverse linguistic backgrounds.

    Source code


  • "Hardware Implementation of a Speech Intelligibility Enhancement Algorithm," K. Shimanuki (2026).
    The main objective of this research is to develop a system that improves speech intelligibility in noisy environments in real-time. It is often difficult to understand speech in environments such as cafeterias and train stations. While simply increasing the sound pressure level (often called "volume") can be an effective way to improve intelligibility, this creates a risk of hearing damage. Therefore, this study focuses on developing a real-time hardware-based signal processing system that enhances speech intelligibility without need of excessive amplification. Furthermore, this study extends previous research by addressing more realistic scenarios involving reverberant environments. The system's performance was evaluated through a subjective listening experiment measuring keyword transcription accuracy. This experiment was conducted with 12 Japanese students. The results indicated a statistically significant improvement in speech intelligibility, achieving an odds ratio of 1.25 compared to the plain condition.

    Source code


  • “Extending markdown to embed pure-data programs (MDPD),” R. Okuda (2025).
    When people write documents on computers, they usually use a simple writing template called Markdown. But if they want to share audio in their writing, it is hard to do so. Users need to open different programs to play sounds, and readers need to download and open audio files separately. This project aims to make it easier for teachers and students who work with audio to play sound demonstrations directly in their Markdown documents. We chose tools such as Node.js and WebPD to build a special program that combines Markdown with Pure-data, a simple way to create and manipulate audio on a computer. The resulting tool lets you view both text and interactive audio in a web browser.

    Source code


  • “Hardware implementation in real-time to improve speech intelligibility,” H. Tanii (2025).
    This study aims to develop a system that improves speech intelligibility in noisy environments in real-time. It is important to improve speech intelligibility to understand announcements in venues such as cafeterias, airports, etc. Previous research has shown that harmonic distortion and dynamic compression are useful for improving speech intelligibility in noisy environments.
    However, the aforementioned research has not been developed for real-time processing. Therefore, in this thesis, we propose a real-time implementation based on a Teensy micro-controller that can apply effects processing to audio in real-time.
    An experiment was conducted with 16 participants, and their accuracy in understanding the listened speech was evaluated. The results indicate no difference between unprocessed speech (plain) and speech treated with real-time implementation. We believe that this could be caused by some limitations in implementing the effect of a pre-emphasis filter.

    Source code


  • “Slight spatial separation affects the psychoacoustic roughness,” K. Suzuki (2025).
    Sound roughness is an important psychoacoustic property that affects auditory perception in a variety of settings, including the evaluation of music and noise. Previous studies have shown that spatial separation of sound sources affects the perception of roughness, but existing roughness models focus primarily on single tones and do not consider binaural spatial effects.
    In this study, we investigated the effect of slight spatial separation on the perception of roughness. Pairs of amplitude-modulated and beating tones were played from spatially separated speakers, and participants adjusted the modulation index to match the perceived roughness. As a result, perceived roughness was reduced even with slight angular separation.
    These results emphasize the need to extend the roughness model to incorporate binaural processing and spatial effects. Future work should investigate roughness perception for sources placed laterally and backward to further refine the binaural roughness model.

    Source code


  • “Perception of haptic missing fundamental,” D. Komagata (2025).
    The present the study of the possibility of missing fundamentals being included in tactile vibrations. This research investigates whether the missing fundamental could also be perceived through touch and whether it exhibits similar characteristics. In the experiments, we asked participants to match the frequencies of haptic stimuli to that of auditory ones whose fundamental was missing. As a result of the experiments, people can perceive the missing fundamental when there is no fundamental frequency. However, we could not measure accurately because the worst SNR value was $7.5$dB. To measure more accurately, we have to use the actuator which is more constant vibration acceleration and do not generate sound leakage.

    Source code


  • “Digital Audio Effect for Simulating Perceived Sound of One's Voice,” Yoshiki Sakai (2024).
    Voice confrontation is a known issue in psychology, where one feels discomfort towards the recorded voice of themselves. Past research have shown the effectiveness of audio filters created through data collection methods involving multiple participants. However, the time and equipment required for acquiring necessary data made the voice conversion technology inaccessible from na\"ive persons. Thus, through this research, we propose the use of non-personalized Head-Related Impulse Responses, obtained from a Head and Torso Simulator, a binaural microphone, and a simulated spherical head, to simulate the perceived sound of one’s voice. The real-time voice conversion filter is implemented as Virtual Studio Technology plugin, through the means of a Finite Impulse Response filter.

    Source code


  • "Implementing real-time risset rhythms," R. Kiuchi (2024).
    The purpose of this paper is to allow modifications of Risset rhythms in real-time at actual concerts. Risset rhythms is an extension of Risset and Shepard tones (an illusion of ever- ascending/descending pitch) to tempo. It is an auditory illusion in which rhythm and tempo seem to rise or fall infinitely. Currently, it has not been utilized effectively in actual concerts or performances. As a method, we worked on the development of Virtual Studio Technology (VST) plug-in for Risset rhythmization using MATLAB’s appdesigner and Pure Data. As a result, we successfully developed a drum machine GUI and the VST plug-in to generate Risset rhythms. Since this plug-in is readily available in Digital Audio Workstation (DAW) applications, it can provide a new option for future concerts and other performances.

    Source code


  • "Spectral delay as digital audio effect," T. Odaira (2024).
    This project presents new Digital Audio Effects (DAFX) which provide users with an interactive sound experience. This project is inspired by the recent discovery of the speed of sound in the acoustic environment on Mars. The program is finally deployed as a Virtual Studio Technology (VST) plugin to enable a user to control/modify the sound experience in real-time.

    Source code


  • "Visualization of three-dimensional auditory trajectories in virtual reality," Y. Ohuchi (2024).
    Our research is about visualization of three-dimensional auditory trajectory in virtual reality. The research method involved developing a tool using Unity to visualize three-dimensional auditory trajectories in using 3D-goggles or a computer screen. It was created based on common equipment found in a laboratory that conducts virtual reality research. Auditory trajectories are color-coded by subject, making it easier to compare trajectories. This research shows that visualizing auditory trajectories eases understanding of three-dimensional auditory trajectory data compared to analyzing them in two dimensions such as tables or 2D-plots.

    Source code


  • "Data Pattern Sonification in Python and Faust," Y. Sato (2022).
    This research presents a tool for the exploration of data patterns via auditory displays. The proof-of-concept software application relies on state-of-the-art data mining techniques (programmed in Python) to find patterns in big data and sonify them using a Frequency Modulation (FM) synthesizer created in Faust. To illustrate the capabilities of the proposed system we sonified atmospheric pollution patterns from a public database.

    Source code


  • "Implementation of an Augmented-Virtuality classroom," D. Sato (2022).
    We developed a proof-of-concept software application suitable for virtual classrooms. Common online discussion tools usually mix participant voices to mono- or stereophonic sound and allow sharing only one material at once. The application proposed in this research features voice memo recording, multiple image sharing, zooming, web camera projection, and spatial voice chatting; improving some of the shortcomings identified in traditional virtual conference tools. At most ten participants can connect to the same room in this application which is demonstrated using a video recordings.

    Source code


  • "Implementation of Voice Directivity in Speech Spatialization for Virtual Environments," Y. Sakurai (2022).
    We aim to create a sound environment similar to real settings for Virtual Reality (VR) conferences. Humans can sense the location of a speaking person and where she or he is facing (orientation). However, when listening to speech through headphones, it is difficult to sense those unless special Head-Related Transfer Functions (HRTFs) are used. To ease the perception of speaker orientation in VR, we recorded HRTFs and created filters that can be applied in real-time. These filter are stored according to the ``Spatially Oriented Format for Acoustics'' (SOFA) standard. In addition, we created a proof-of-concept application to manipulate sound sources with Pure-data.

    Source code


  • "Loudspeaker Sound Spatialization via Precedence Effect," Y. Minakawa (2022).
    The precedence effect is a useful phenomenon that eases the perception of sound direction in reverberant environments. We created an audio spatialization program based on this effect that allows sound spatialization via loudspeaker arrays. The location of the loudspeakers can be arbitrarily set, their convex hull is used as an inner room with imaginary soundproof walls so sound can only be projected from the ``windows,'' i.e., loudspeakers. Our system allow the spatialization of audio sources outside the inner room, but inside a larger outer room with walls used to compute first reflections, also added to the spatialization. Our proof-of-concept application successfully produces the spatialization of sounds as demonstrated in a video.

    Source code


  • "Real-time Binaural Audio in Pure-data Based on Head-Related Transfer Functions," K. Fujita (2022).
    We introduce a real-time HRTF-based spatializer system in Pure-data using a database that contains compressed HRTF information. The system increased the density of filters, requiring less storage than previous systems. In addition, this system also reduces CPU usage about $25$\,\% to $50$\,\% relative to a previous similar system.

    Source code


  • "Real-time Manipulation of Spectral Tilt of Audio Signals," S. Yoshida (2022).
    Measuring spectral tilt is one of the most important metrics in analyzing audio. However, it can also be used to modify the audio signal by manipulating the spectral tilt. The goal of this study is to implement the application of the manipulation of spectral tilt of audio signals in real-time. This application can manipulate the spectral tilt of the input monaural signal and output these. We implemented an original object by using Pure-data. And, we also implemented two applications with it and evaluated them.

    Source code


  • “ Spatially Oriented Format for Acoustics (SOFA) reader object for Pure-data ” S. Fujisawa (2021)
    My research is about writing objects in Pure-data that can be used for sound spatialization by reading and using databases specified according to the Spatially Oriented Format for Acoustics (SOFA) standard.

    Source code


  • “ Auralization of Three-dimensional Cellular Automata” Y. Kariyado (2021)
    This thesis presents an implementation of three-dimentional cellular automata and auralization tool. The proof-of-concept of this project enables the creation of the spatialized sounds associated with cells in a dimensional grid. This proposed tool promotes exploring and storing new patterns of cellular automata, and can also be used for generating music in real time.

    Source code


  • “ Subjective Evaluation System for Audio in VR” M. Tsukamoto (2021)
    This research aims at developing a subjective evaluation system for audio in VR. If the angle and orientation of the listener's head changes, the apparent position of the audio source in VR and the 3D sound in VR will not work properly. To solve this problem, we have developed the current system. In addition, we sought to simplify the tasks of the subjects and the experimenters. As a result of this project, the subjective accuracy of the experiments is expected to improve, and in addition, the experimental efficiency is improved.

    Source code


  • “ Generalized Transaural Spatialization for Pure-data” R. Kakiki (2021)
    A transaural system allows to reproduce binaural sounds via loudspeakers. In our lab, we have created a Pure-data object that implements a transaural system for the speakers in front of the listener. The problem with the previous program is that when the listener moves, the spatialization is destroyed. This study aims to improve the spatialization of transaural objects by selecting two active loudspeakers depending on the azimuth of the listener.

    Source code


  • “ Virtual Reality Navigation system using music for optimal traffic speed” Y. Akahori (2021)
    The conventional traffic system forced us stop and go system. That system is the best way in safety, but some people are tired by waiting traffic lights in system. On the other hand, there is also an attempt to create a smooth traffic network by controlling multiple traffic lights, known as a green wave. We created a system to find the timing of the green wave and other factors that would induce speed for the user. The goal is also to use music as a medium for speed navigation at that time, so as not to obstruct the user’s vision.

    Source code


  • “Synthesis of binaural signals with arbitary trajectories” R. Agata (2020)
    The purpose of this study is to reproduce recorded sounds with binaural sound. To do it, we made an interface that can record an instrument’s movement. The interface can be edited easily because of VR can visualize movement. After recording the trajectories, the interface renders a binaural audio file. This interface can record the trajectories of up to five instruments.

    Source code


  • “Improving bass perception with vibration motors” Y. Suzuki (2020)
    When listening to music, people often use headphones or earphones and turn the sound level up. Because of the size of the diaphragms, it is difficult to hear the bass as heard for example in a concert. So, this could cause serious problems included damaging the hearing system. In this research, we explore a way to reduce this problem by feeling the bass from vibration motors. We considered to add the bass amplification by vibration to the music we usually hear. We thought this way would be the solution in which listeners do not have to turn the sound level up.

    Source code


  • “Virtual Reality Theremin” Y. Igarashi (2020)
    In this project, we create Virtual Theremin by using APIs that Unity, Pure-data, HRIRu. The theremin is the first electronic musical instrument in the world. There are not many contents of VR musical instruments. If the VR musical instrument becomes popular, there is no need to go to the studio to practice and mental image training is possible to perform in front of many audiences in VR. We try to explore the potential of the VR instrument by creating Virtual Theremin.

    Source code


  • “Floor reflection effect on elevation perception of sounds,” D. Hasegawa (2019).
    The aim of this study is to improve accuracy of sound elevation perception. Compared to the subjective accuracy of artificial sounds above a listener’s ears, that achieved for artificial sounds below is usually reported to be inferior. Because of that, reflection of sound and no reflection of sound were recorded in an anechoic chamber. Recorded signals were processed to eliminate the effect of the recording apparatus (microphone, loudspeaker, etc.). In this experiment, four methods (direct, delay, reflection, and mixed HRIR) were used. Participants listened to sound processed with these methods and judge their elevation. As a result, accuracy to distinguish sounds from below ears improved by using HRIRs including reflection from the floor.

  • “Improving Speech Localization in Virtual Reality,” N. Miyauchi (2019).
    The purpose of this study is to make a database of HRIRs (Head Related Impulse Responses) with mouth radiation. In this thesis, we collected impulse responses by using two HATS (Head and Torso Simulator) in an anechoic chamber. We used equalizing filters to create sound image of the direction by making IR convoluted of mono sound. An experiment was performed asking how people perceive a directed sound source. The average error is 39.2°. The results indicate that sound perception had an effect on by using two HATS at listener azimuth 0°. This project is in collaboration with Ruri Moriyama.

  • “Improving speech localization in VR,” R. Moriyama (2019).
    The aim of this study is to measure voice directivity using two HATS (Head and Torso Simulators), one as a listener and the other as a speaker. We often know the location of a speaker (right, left, back, front, etc.) in virtual reality (VR). But, it is difficult to understand the orientation of a speaker in these environments because HRIRs (Head-Related Impulse Responses) have no voice directivity pattern built into them. On the contrary in the real world, we can understand the location and orientation of a speaker with good accuracy even when he or she cannot be seen. The HATS we used as speaker featured a mouth simulator with a directivity pattern similar to at found in humans. We used diffuse-field equalization before creating spatialized sounds. An experiment was carried out to verify how people perceive spatialized sounds comparing with a normal database and our database. This study’s result indicates that it was more difficult for people to perceive speaker’s orientation correctly from the side in VR scenes than from the front. This is a collaboration project with Ms. Miyauchi because of its large scope.

  • “Replacing vocalic part with Shepard-tone like spectrum in real-time,” S. Hirata (2019).
    The aim of study is exploring aesthetic possibilities of sung voice with inharmonic instruments and real-time vocalic detection for future applications. In this thesis, we used a Pure-data program to replace vocalic part with Shepard-tone like spectrum in real-time (same pitch and same reported duration). An experiment was performed investigating how psychoacoustic roughness of voice with inharmonic instrument minimize. Participants were asked to estimate roughness they listened of four types of sound sources with different ratios. The results indicates that roughness of voice with inharmonic timbre is not minimized by using Pure-data program.

  • “Platform for comparing sound spatialization methods in virtual reality,” A. Uemura (2018).
    This research aimed at building an acoustic environment for Virtual Reality (VR) spatialization of sound. To test the built VR program, a subjective experiment was conducted to compare the judgements of virtual sound images processed with two methods: Default spatializ tion in Unity and using HRTF convolution. The results indicated that the proposed environment could be used for other method comparisons in the future.

  • “Perception of spatialized Risset tones,” N. Fukasawa (2018).
    This research aimed at building an acoustic environment for Virtual Reality (VR) spatialization of sound. To test the built VR program, a subjective experiment was conducted to compare the judgements of virtual sound images processed with two methods: Default spatializ tion in Unity and using HRTF convolution. The results indicated that the proposed environment could be used for other method comparisons in the future.

  • “Improving sound perception in elevation using single layer loudspeaker array display,” Y. Suzuki (2018).
    The purpose of this study is to improve the localization in elevation of sound sources coming from a single-layer loudspeaker array. In this thesis, we used a method using equalizing filters to create elevated sound images. The experiment was performed eliciting how people perceive the elevated sound source. They were asked the direction they perceive of the sound using elevation and azimuth angle. The results indicate that sound perception was improved by using a loud- speaker grouping method.

  • “Computer-assisted singing experience,” M. Ishihara (2018).
    This research aimed at building an acoustic environment for Virtual Reality (VR) spatialization of sound. To test the built VR program, a subjective experiment was conducted to compare the judgements of virtual sound images processed with two methods: Default spatialization in Unity and using HRTF convolution. The results indicated that the proposed environment could be used for other method comparisons in the future.

  • “Assisting System for Grocery Shopping Navigation and Product Recommendation,” S. Saito (2018).
    We present a system for grocery shopping by recommending products according to those currently in a shopping basket, and supporting users in finding their way to the location of recommended item in store.
    Our goal is supporting users in finding their way from their current position to the location of a determined item. And suggesting items to user that analysis by big data.

  • “Implementation of a transaural system in Pure-data,” T. Ninagawa (2017).
    Transaural audio is a method used to deliver binaural signals to a listener using regular stereo loudspeakers. This paper discusses the implementation of a transaural audio filter in Pure-data. Pure-data is a visual programming language used to process and generate sound, video, 2D/3D graphics, and interfaces sensors, input devices, etc.

  • “Quantifying the benefits of bimodal navigation systems,” T. Takahashi (2015).

  • “Loudness perception with headphone and vibration,” Y. Ito (2015).
    When listening to music, people often listen at high levels because it is difficult to hear the bass. This causes a variety of problems. The number of persons who listen to music everyday is increasing, and so it is hearing impairments. This research aims to find ways to reduce these problems by exploring the energy of the bass from music. We used vibration motors move amplify the low frequency energy found in the bass. It is assumed that the listener was difficult to hear the bass and that is one a reason to listen to music at high levels. In view of this, we considered to add bass amplification by vibration to the music that is routinely heard. It is expected that users do not have to increase the music level.

  • “A study of ultrasound encoding and decoding based on steganography,” R. Igarashi (2015).
    This paper describes a method of the information communication mean using the steganography that is one of the information hiding technologies. This technique uses a sound, but does not use the conventional electric wave that is used for a cell-phone and the Internet. This technique suggested in this report conceals information in a high frequency band beyond a human being’s audible range. It makes possible to transmit and receive information without being recognized by a third party. Therefore, the letter is transmitted by the use of suggestion technique, and inspect its effective- ness with decoding reverse processing in the receiver.

  • “Steganography in stereo signals using phase changes,” S. Hoshi (2014).
    The present study was undertaken in order to embed information into the right channel of a stereo audio file and to create its corresponding detection system. In this research, we focus on a technique of phase modulation to watermark information. Phase modulation was implemented with all-pass filters. Information to be embedded was the binary data corresponding to an ASCII coded message. This simple approach shows that it is possible to embed information this way with applications in in-door localization, security, etc.

  • “Navigating in virtual worlds using smartphones: Reflecting real world motion in virtual environments,” Y. Chiba (2013).
    “Freedom” is a project that lets navigate a virtual enviroment using sensors in smartphones. Users can control 3D animation created in Alice, software for creating and rendering 3D animation by Java programming language. “Freedom” is a new way of controlling 3D animation intuitively and impressing users with realistic experience.
    Note: In collaboration with Prof. Cohen

  • “`Machi-Beacon’: Spatial Sound For Mobile Navigation System,” W. Sanuki (2013). This research explores the development of mobile navigation systems using spatial sound. A combination of spatial sound and geographic information allows mobile device users to perform auditory localization tasks while their eyes, hands, and attention are occupied. We created a spatial sound and simple navigation system using Unity. This navigation system informs user orientation to a goal and distance between present place and that point. The system runs on iOS and Android OS. Preliminary results indicate that a combination of spatial sound GPS and GIS can be used for navigation.

    This research was presented at the “Aizu Industry IT Technology” contest where it won an Encouragement Prize.
    Note: In collaboration with Prof. Cohen


  • “ Implementing an A-weighting filter as an external object in Pure-data,” K. Sakui (2013).
    This research aimed the construction of A-weighting filter as an external object in Pd. C language was used for building the external. The same processing as actual A-weighting and created external object to perform, and the result was compared using a graph. Similar result were obtained in the middle frequency range. But, for low frequencies (under 100 Hz), values differed noticeably.
    Note: In collaboration with Prof. Cohen