Journal articles

  1. Kimi Akita, Shigeto Kawahara, and Julián Villegas. Smiles can be conveyed to listeners via speech signals: A case study with Japanese labiodentals. Linguistics Vanguard, Aug 2026. DOI: 10.1515/lingvan-2026-0011.

    One long-standing question in speech communication research concerns the extent to which an emotional state is encoded in the acoustic signal and conveyed to the listeners. To address this question, we deployed a recent finding in Japanese, in which speakers replace bilabial sounds (e.g., [b]) with labiodental sounds (e.g., [v]) when speaking with a smile. Experiment 1 used stimuli recorded by two professional voice actors under four conditions crossing two factors, Voice Quality (Smile vs. Non-Smile) and Articulation (Labiodental vs. Bilabial). The results confirm that labiodentals function as a reliable cue for smiling, even in the absence of visual information. Experiment 2 controlled for the effects of F0 and amplitude of the stimuli, and we still observed an effect of labiodentals. We conclude (1) that smiles can be conveyed via speech signals to the listeners without any visual cues and (2) that in the case of Japanese, this is achieved via the use of labiodentals.

  2. Camilo Arévalo and Julián Villegas. Spatial upsampling of head-related impulse responses via elevation-wise autoencoders. IEEE Open J. of Signal Processing (OJSP), Sep. 2025. DOI: 10.1109/OJSP.2025.3613209.

    A method for performing spatial upsampling of Head-Related Impulse Responses (HRIRs) from sparse measurements is introduced. Based on an elevation-wise autoencoder design, we present two variants: one that performs progressive reconstructions with feed-forward connections from higher to lower elevations, and another that excludes these connections. The variants were evaluated in terms of the errors in interaural time and level differences, as well as the spectral distortion in the ipsilateral and contralateral ears. The additional complexity introduced by the variant with feed-forward connections does not always translate into accuracy gains, making the simpler variant preferable for efficiency. Furthermore, the error decreased as the number of measurements used for upsampling increased. However, even with only three measurements, the proposed method surpasses the accuracy of using an average HRIR across listeners while achieving differences similar to or smaller than those that could be perceptible to listeners

  3. Julián Villegas. A new method for speech enhancement based on harmonic distortion. IEEE Transactions on Audio, Speech and Language Processing, pages 3826–3837, Sep. 2025. DOI: 10.1109/TASLPRO.2025.3608964.

    A near-end listening enhancement (NELE) method that combines harmonic (non-linear) distortion with linear distortion for improving speech intelligibility in noisy conditions without altering the signal-to-noise ratio (SNR) is proposed. Parameters were optimized across noise types, SNRs, and reverberation using a Genetic Algorithm with Speech Intelligibility in Bits (SIIB) as the loss function. Subjective experiments confirmed that, on average, the odds of correctly recognizing a word treated with the proposed method were 18.8 higher relative to plain speech when presented at −9 dB SNR, averaged over the results of Cafeteria and Speech-Shaped Noise (SSN). In Cafeteria noise, the odds of correct recognition were 16.4 times higher compared to plain speech, while in SSN they were about 8.4 times higher, averaged across SNRs ranging from −9 to −3 dB. Objective evaluation further confirmed its superiority to alternative approaches in reverberant maskers. These results demonstrate that combining harmonic and linear distortions can substantially improve intelligibility in adverse listening conditions.

  4. Timur Jaganov, John Blake, Julián Villegas, and Nicholas Carr. Large language model-driven dynamic assessment of grammatical accuracy in English language learner writing. IEEE Access, 13:151538–151550, Aug. 2025. DOI: 10.1109/ACCESS.2025.3603191.

    This study investigates the potential for Large Language Models (LLMs) to scale-up Dynamic Assessment (DA). To facilitate such an investigation, we first developed DynaWrite—a modular, microservices-based grammatical tutoring application which supports multiple LLMs to generate dynamic feedback to learners of English. Initial testing of 21 LLMs, revealed GPT-4o and neural chat to have the most potential to scale-up DA in the language learning classroom. Further testing of these two candidates found both models performed similarly in their ability to accurately identify grammatical errors in user sentences. However, GPT-4o consistently outperformed neural chat in the quality of its DA by generating clear, consistent, and progressively explicit hints. Real-time responsiveness and system stability were also confirmed through detailed performance testing, with GPT-4o exhibiting sufficient speed and stability. This study shows that LLMs can be used to scale-up dynamic assessment and thus enable dynamic assessment to be delivered to larger groups than possible in traditional teacher-learner settings.

  5. Irene de la Cruz Pavía, Julián Villegas, Caroline Nallet, Kyoji Iwamoto, Ramón Guevara, and Judit Gervain. The role of syllabic rhythm in speech perception across languages. Scientific Reports, 15(1):24494, Jul. 2025. DOI: 10.1038/s41598-025-07053-y.

    The insertion of silences at regular intervals restores the intelligibility of English utterances that have been accelerated beyond comprehension, as long as the duration of the resulting speech-silence chunks falls within the theta rhythm of natural speech, i.e. the temporal modulation associated to the syllabic rate. We test whether such a rhythmic strategy works in languages rhythmically different from English, a stress-timed language. Thus, we assess whether comprehension of time-compressed Semantically Unpredictable Sentences (SUS) is restored in the syllable-timed language French and the mora-timed language Japanese, when silences re-establishing theta rhythm are inserted. Restoring the theta rhythm also improved intelligibility in French, but not in Japanese, in which best performance was instead achieved at faster rhythms, which suggests that modulation at the rate of a language’s basic rhythmic unit plays a key role in understanding speech. In a second experiment, French speakers listened to SUS with speech-silence chunks adapted to the range of the temporal modulations of the delta, gamma, and high gamma rhythms, which correspond to the rate of prosodic phrases, phonemes, and subsegmental features, respectively. Unlike the theta rhythm, we found no restorative effects, providing further evidence for the special status of the theta rhythm in speech comprehension.

  6. Jeremy Perkins, Seunghun J. Lee, and Julián Villegas. The interaction between tone and duration in Du’an Zhuang. J. of the Int. Phonetic Assoc., Feb. 2025. DOI: 10.1017/S0025100324000239.

    This research presents a phonetic production study of the tone system of an understudied language, Du’an Zhuang. Six native speakers were recorded in Guangxi, China and acoustic analyses of f0 and duration reaffirmed previous impressionistic results that found f0 contour to be the primary cue for tone. There are six contrastive tones among unchecked syllables, reduced to four tones in checked syllables. Only two of the four checked tones had matching allotones among the unchecked tones, differing from previous findings where all four checked tones had matching unchecked allotones. Rhyme duration did not differ among the tones, except that the tones with falling f0 contours had shorter durations. Finally, this study also investigated how phonological vowel length contrasts were preserved given that rhyme durations are held constant, finding that sonorant coda durations compensated for vowel length: Sonorant codas were lengthened following short vowels and were shortened following long vowels

  7. Kazuki Fujita and Julián Villegas. Development of ad hoc loudspeaker arrays using smart devices. IEEE Communications Magazine, Dec. 2024. DOI: 10.1109/MCOM.001.2400155.

    Traditional loudspeaker arrays are effective in creating immersive auditory experiences but face real-time challenges related to adaptability, particularly regarding the number and location of speakers. Although smart devices have the potential to address these limitations, there is a lack of research exploring their use, especially in indoor settings. Our study addresses this gap by proposing an ad hoc loudspeaker array system that leverages smart devices for real-time audio spatialization. The system employs a modified “room within a room” spatialization method, dynamically defining the listening area using the convex hull formed by the loudspeakers. Ultra-wideband (UWB) technology is utilized for precise indoor location, allowing speakers to move in the horizontal plane while maintaining accurate audio spatialization. Developed as an iOS application, the system was inherently limited in the number of interconnecting devices, and was subjectively evaluated using both moving and stationary stimuli and loudspeakers. Results consistently indicated positive ratings for moving stimuli, regardless of speaker motion, while stationary stimuli received lower ratings. This suggests potential areas for improvement, such as increasing the number of devices allowed to connect. These findings demonstrate the feasibility of using commercially available smart devices to create flexible loudspeaker arrays, with implications for the development of more adaptable and user-friendly spatial audio systems.

  8. James Pinkl, Julián Villegas, and Michael Cohen. Multimodal drumming education tool in mixed reality. Multimodal Technol. and Interact., 8(8), Aug. 2024. DOI: 10.3390/mti8080070.

    1st-person VR-based Action Observation research has thus far yielded both positive and negative findings in studies observing such tools’ potential to teach motor skills. In this contribution, a 1st-person mixed reality-based tool with multimodal cues designed to teach rudimental and polyrhythmic drumming was developed and tested in a 20-subject study. When compared against a control group practicing via video demonstrations, results showed increased rhythmic accuracy across four exercises. Specifically, a difference of 239 ms (z-ratio=3.520,p¡0.001) was found between the timing errors of subjects who practiced with our multimodal mixed reality development compared to subjects who practiced with video, demonstrating potential of such affordances. This research contributes to ongoing work in the fields of Action Observation and Mixed Reality, providing evidence that Action Observation techniques can be an effective practice method for drumming.

  9. Edward Ly and Julián Villegas. Cartesian genetic programming parameterization in the context of audio synthesis. IEEE Signal Process. Letters, 30:1077–1081, Aug. 2023. DOI: 10.1109/LSP.2023.3304198.

    This article presents an evaluation of the effects of elitism, recurrence probability, and prior knowledge on the fitness achieved by Cartesian Genetic Programming (CGP) in the context of DSP audio synthesis. Prior knowledge was introduced using a probabilistic learning method where the distribution of nodes in the expected solutions was used to generate and mutate new individuals. Best results were obtained with traditional elitist selection, no recurrence, and when prior knowledge was used for node initialization and mutation. These results suggest that the apparent benefits of recurrence in CGP are context-dependent, and that selecting nodes from a uniform distribution is not always optimal.

  10. Julián Villegas, Seunghun J. Lee, Jeremy Perkins, and Konstantin Markov. Psychoacoustic features explain creakiness classifications made by naive and non-naive listeners. Speech Comm., 147:74–81, Jan. 2023. DOI: 10.1016/j.specom.2023.01.006.

    We compared the classifications of words between creaky and non-creaky made by two groups of listeners whose native languages differ on whether creakiness is used for phonemic contrast. Japanese (naive group) and Vietnamese (non-naive group) listeners classified words produced by a single speaker of Du’an Zhuang, an unintelligible language for both groups that shares some tonal similarities with Hanoi Vietnamese. We found little differences between the accuracy achieved by the two groups, and that the non-naive group was more likely to rate words as creaky relative to the naive group. In addition, there seems to be no benefit of background linguistic knowledge on the accuracy with which the non-naive group classified words. The sensitivity to creakiness observed in classifications made by experts (inspecting the waveforms and spectrograms in addition to listening to the words) and made by a machine learning algorithm based on these kind of classifications was unmatched by both cohorts. This result challenges the perceptual validity of such refined classifications commonly used in phonation studies. In addition, we found a positive association between psychoacoustic roughness and the probability of a word to be judged as creaky. We also found a positive association of loudness with creaky judgments, whereas pitch was negatively associated. We found no evidence of sharpness and creaky association. Finally, the accuracy to predict subjective creakiness via a recurrent neural network classifier was best when the traces of all the considered psychoacoustic features were included as predictors.

  11. Seunghun J. Lee, Julián Villegas, and Mira Oh. The Non-Coalescence of /h/ and Incomplete Neutralization in South Jeolla Korean. Language and Speech, 66(2), Aug 2022. DOI: 10.1177/00238309221116130.

    This study examines articulatory and acoustic data in order to investigate the non-coalescence of /h/ in South Jeolla. Seoul Korean speakers produce /pap/ “rice” followed by /hana/ “one” as [pa.pha.na] with the coalescence of /p/ and /h/; this is called an aspiration merger. In South Jeolla Korean, this merger may be blocked, as in cases where speakers produce /pap+hana/ as [pa.ba.na]. Electroglottographic (EGG) data indicate the existence of two groups of South Jeolla speakers: one that merges the plosive and /h/ (the merger group), and the other with the canonical South Jeolla Korean pronunciation that does not merge the two consonants (the non-merger group). The production of non-coalesced lenis stops in the non-merger group is phonetically comparable with an underlying lenis stop produced by both of the groups. However, in the non-merger group, the open quotient (OQ) of a vowel following a non-coalesced lenis stop is higher (breathier) than that of an underlying lenis stop. Spectral tilt results display a similarly increased breathiness when the vowel follows a non-coalesced lenis stop. As for the non-merger group of South Jeolla, we argue that speakers display incomplete neutralization such that the non-merger group produces two types of voiced lenis stops differing in the phonation of the following vowel. These findings suggest that previous phonological analyses that posit the /h/-deletion in the non-merger group of South Jeolla Korean need to be revisited.

  12. Camilo Arévalo, Yuta Kariyado, and Julián Villegas. Virtual reality tool for exploration of three-dimensional cellular automata. Electronics, 11(3), Feb 2022. DOI: 10.3390/electronics11030497.

    We present a Virtual Reality (VR) tool for exploration of three-dimensional cellular automata. In addition to the traditional visual representation offered by other implementations, this tool allows users to aurally render the active (alive) cells of an automaton in sequence along one axis or simultaneously creating melodic and harmonic textures while preserving in all cases the relative locations of these cells to the user. The audio spatialization method created for this research can render the maximum number of audio sources specified by the underlying software (255) without audio dropouts. The accuracy of the achieved spatialization is unrivaled since it is based on actual distance measurements as opposed to coarse distance approximations used by other spatialization methods. A subjective evaluation (effectively, self-reported measurements) of our system (n = 30) indicated no significant differences in user experience or intrinsic motivation between VR and traditional desktop versions (PC). However, participants in the PC-group explored more of the universe than the VR-group. This difference is likely to be caused by the familiarity of our cohort with PC-based games.

  13. Julián Villegas, Naoki Fukasawa, and Camilo Arevalo. The presence of a floor improves subjective elevation accuracy of binaural stimuli created with non-individualized head-related impulse responses. J. Audio Engineering Society, 69(11):849–859, Nov. 2021. DOI: 10.17743/jaes.2021.0045.

    We report the effect of the presence of floor on elevation estimation of audio spatialized with non-individualized Head-Related Impulse Responses (HRIRs). The results of two experiments (n = 21 and n = 39) suggest that using HRIRs captured when a floor simulator was present improved assessors’ accuracy when judging elevation in the sagittal and coronal planes, especially at high elevations. Such improvements were not observed when signals delayed according to their computed first reflection were mixed with signals convolved with anechoic HRIRs. These findings suggest that capturing non-individualized HRIRs in hemi-anechoic rooms could improve accuracy of audio spatialization in virtual environments.

  14. Julián Villegas, Jeremy Perkins, and Ian Wilson. Effects of task and language nativeness on the Lombard effect and on its onset and offset timing. J. Acoust. Soc. Am., 149(3):1855–1865, Mar 2021. DOI: 10.1121/10.0003772.

    This study focuses on the differences in speech sound pressure level (here called speech loudness) of Lombard speech (i.e., speech produced in the presence of an energetic masker) associated with different tasks and language nativeness. Vocalizations were produced by native speakers of Japanese with normal hearing and limited English proficiency while performing four tasks: dialog, a competitive game (both communicative), soliloquy, and text passage reading (non- communicative). Relative to the native language (L1), larger loudness increments were observed in the game and text reading when performed in the second language (L2). Communicative tasks yielded louder vocalizations and larger increments of speech loudness than non-communicative tasks, regardless of the spoken language. The period in which speakers increased their loudness after the onset of the masker was about fourfold longer than the time in which they decreased it after the offset of the masker. Results suggest that when relying on acoustic signals, speakers use similar vocalization strategies in L1 and L2, and these depend on the complexity of the task, the need for accurate pronunciation, and the presence of a listener. Results also suggest that speakers use different strategies depending on the onset or offset of an energetic masker.

  15. Edward Ly and Julián Villegas. Generating artificial reverberation via genetic algorithms for real-time applications. Entropy, 22(11):1309, Nov. 2020. DOI: 10.3390/e22111309.

    We introduce a Virtual Studio Technology (VST) 2 audio effect plugin that performs convolution reverb using synthetic Room Impulse Responses (RIRs) generated via a Genetic Algorithm (GA). The parameters of the plugin include some of those defined under the ISO 3382-1 standard (e.g., reverberation time, early decay time, and clarity), which are used to determine the fitness values of potential RIRs so that the user has some control over the shape of the resulting RIRs. In the GA, these RIRs are initially generated via a custom Gaussian noise method, and then evolve via truncation selection, random weighted average crossover, and mutation via Gaussian multiplication in order to produce RIRs that resemble real-world, recorded ones. Binaural Room Impulse Responses (BRIRs) can also be generated by assigning two different RIRs to the left and right stereo channels. With the proposed audio effect, new RIRs that represent virtual rooms, some of which may even be impossible to replicate in the physical world, can be generated and stored. Objective evaluation of the GA shows that contradictory combinations of parameter values will produce RIRs with low fitness. Additionally, through subjective evaluation, it was determined that RIRs generated by the GA were still perceptually distinguishable from similar real-world RIRs, but the perceptual differences were reduced when longer execution times were used for generating the RIRs or the unprocessed audio signals were comprised of only speech.

  16. Julián Villegas, Konstantin Markov, Jeremy Perkins, and Seunghun J. Lee. Prediction of creaky speech by recurrent neural networks using psychoacoustic roughness. IEEE J. of Selected Topics in Signal Processing, 14(2):355–366, Feb. 2020. DOI: 10.1109/JSTSP.2019.2949422.

    The use of a psychoacoustic roughness model as a predictor of creaky voice is reported. We found that the roughness temporal profile of vocalic segments can predict the presence of creakiness in speech. Using a simple bi-directional Recurrent Neural Network (RNN), we were able to predict the presence of creakiness in vocalic segments from only roughness traces with an accuracy similar to that obtained with RNNs trained on at least 12-dimensional input data (including amplitude difference between the first two harmonics, residual peak prominence, etc.). Training RNNs with the combination of roughness and multidimensional input data improved the performance of the predictor, but not significantly. Likewise, augmenting the dataset by time derivatives of the input features did not improve the predictor’s performance. The proposed roughness-based predictor eases interpretation and comparison of creakiness among corpora and suggests that roughness prediction models could be successfully used for classification of creaky intervals in speech.

  17. Irene de la Cruz Pavía, Gorka Elordieta Alcibar, Julián Villegas, Judit Gervain, and Itziar Laka. Segmental information drives adult bilingual phrase segmentation preference. Int. J. of Bilingual Education and Bilingualism, 25(2):676–695, Jan. 2020. DOI: 10.1080/13670050.2020.1713045.

    In two artificial language learning experiments with four groups of highly proficient Basque-Spanish bilinguals and two groups of Spanish monolinguals, we examine the cues that allow adult listeners to parse new input into phrases. In addition, we investigate which factors lead bilinguals to switch between the segmentation strategies characteristic of their two languages. We show that segmental information drives bilinguals’ choice of a segmentation strategy when presented with an unfamiliar language. The language in which bilinguals are addressed during the study (i.e., the language of context) additionally modulates their segmentation preference, and this language context effect is found in L1Basque bilinguals but does not extend to L1Spanish bilinguals. The cause of this asymmetry is yet to be established. Finally, we show that adult monolinguals disregard statistical cues in favor of unfamiliar segmental information when in conflict. These results evidence that the available phrase segmentation cues are arranged hierarchically.

  18. Julián Villegas. Movement perception of Risset tones presented diotically. Acoustic. Sci. & Tech., 41(1), Jan 2020. DOI: 10.1250/ast.41.430.

    In previous research, we found that when asked to estimate the apparent origin and motion of Risset tones, subjects were more likely to associate them with horizontal than with vertical movements ones. A new experiment was conducted in which a greater freedom of choice for subjective answer was introduced. The present study presents the results of that experiment and aims at replicating the previously found results for unspatialized Risset tones with upward and downward glissandi, finding whether Risset tones with different frequency component separation yield different outcomes, and whether Risset tone opinions differ from those of speech.

  19. Julián Villegas. Improving perceived elevation accuracy in sound reproduced via a loudspeaker ring by means of equalizing filters and side loudspeaker grouping. Acoust. Sci. & Tech., 40(2):127–137, Mar 2019. DOI: 10.1250/ast.40.127.

    A spatialization method for loudspeaker arranged in a ring is presented. The proposed method can be described as an improvement to Vector-Base Amplitude Panning (VBAP) in which a frequency-dependent gain is applied to mitigate the effect of loudspeaker locations on the desired signal. By grouping side loudspeakers it is possible to display elevated sources, which is difficult to achieve otherwise. Objective evaluations showed that the proposed method produces less spectral distortion than traditional VBAP, but degrades to cross-talk cancellation. Experimental results suggest that, despite the differences regarding cross-talk cancellation, the proposed method yields accurate azimuth and elevation estimations of sound sources in anechoic and echoic conditions except when they are located below the ear level.

  20. Julián Villegas and Naoki Fukasawa. Doppler illusion prevails over Pratt effect in Risset tones. Perception, 47(12):1179–1195, 2018. DOI: 10.1177/0301006618807338.

    Changes in frequency such as those found in Risset tones have been associated with moving sound sources in the vertical plane (Pratt effect) and the horizontal plane (Doppler illusion). We investigated the reported origin and motion of unspatialized Risset tones presented monotically and diotically, and Risset tones simulated to be in the sagittal or coronal plane, approaching or receding, from above or horizontally. Independent of the artificial spatialization used (none, spatializing frequency components collectively or individually, elevated or not), upward glissandi were more likely to be judged as approaching than receding, and downward glissandi as receding than approaching, in most cases from the horizon. Glissandi associations with horizontal movements were more common in stimuli simulated on the sagittal plane than in stimuli simulated on the coronal plane. These findings suggest that the Doppler illusion is stronger than the Pratt effect, at least for Risset tones presented over headphones and simulated to be in the sagittal plane. These findings may contribute to better understanding of the association between auditory motion perception and changes in frequency.

  21. Jorge González-Alonso, Julián Villegas, and M.P. García-Mayo. English compound and non-compound processing in bilingual and multilingual speakers: Effects of dominance and sequential multilingualism. Second Language Research, 32(4):503–535, May 2016. DOI 10.1177/0267658316642819.

    This article reports a study which investigated the relative influence of the first and dominant language on L2 and L3 morpho-lexical processing. A lexical decision task compared the responses to English NV-er compounds (e.g. taxi driver) and non-compounds provided by a group of native speakers and three groups of learners at various levels of proficiency in English: L1 English-L2 Spanish sequential bilinguals and two groups of early Spanish-Basque bilinguals with English as their L3. Crucially, the two trilingual groups differed in their first and dominant language (i.e. L1 Spanish-L2 Basque vs. L1 Basque-L2 Spanish). Our materials exploit an (a)symmetry between these languages: while Basque and English pattern together in the basic structure of NV-er compounds, Spanish presents a very different construction. Results show differences in response times that may be ascribable to two factors beyond proficiency: the number of languages spoken by a given participant and the nature of their L1. An exploration of response bias reveals an influence of the participants’ L1 on the processing of NV-er compounds. Our data suggest that morphological information in the nonnative lexicon may extend beyond morphemic structure, that there are costs to additive multilingualism in lexical retrieval, and that most of these effects are attenuated by proficiency.

  22. Julián Villegas. Locating virtual sound sources at arbitrary distances in real-time binaural reproduction. Virtual Reality, 19(3):201–212, Oct 2015. DOI: 10.1007/s10055-015-0278-0.

    A real-time system for sound spatialization via headphones is presented. Conventional headphone spatialization techniques effectively place sources on the surface of a virtual sphere around the listener. In the new system, sources can be spatialized at different distances from a listener by interpolating Head-Related Impulse Responses (HRIRs) measured between 20 and 160cm. These HRIRs are stored in different databases depending on the audio sampling rate. To ease the realtime constraints, users can choose the number of hrir taps used in the convolution, and an alternative interpolation technique (simplex interpolation) was implemented instead of trilinear interpolation. Subjective tests showed that such simplifications yield satisfactory spatialization for some angles and distances.

  23. Martin Cooke, Catherine Mayo, and Julián Villegas. The contribution of durational and spectral changes to the Lombard speech intelligibility benefit. J. Acoust. Soc. Am., 135(2):874–883, Feb 2014. DOI: 10.1121/1.4861342.

    Speech produced in the presence of noise (Lombard speech) is typically more intelligible than speech produced in quiet (plain speech) when presented at the same signal-to-noise ratio, but the factors responsible for the Lombard intelligibility benefit remain poorly understood. Previous studies have demonstrated a clear effect of spectral differences between the two speech styles and a lack of effect of fundamental frequency differences. The current study investigates a possible role for durational differences alongside spectral changes. Listeners identified keywords in sentences manipulated to possess either durational or spectral characteristics of plain or Lombard speech. Durational modifications were produced using linear or nonlinear time warping, while spectral changes were applied at the global utterance level or to individual time frames. Modifications were made to both plain and Lombard speech. No beneficial effects of durational increases were observed in any condition. Lombard sentences spoken at a speech rate substantially slower than their plain counterparts also failed to reveal a durational benefit. Spectral changes to plain speech resulted in large intelligibility gains, although not to the level of Lombard speech. These outcomes suggest that the durational increases seen in Lombard speech have little or no role in the Lombard intelligibility benefit.

  24. Michael Cohen, Rasika Ranaweera, Hayato Ito, Shun Endo, Sascha Holesch, and Julián Villegas. “twin spin”: Steering karaoke (or anything else) with smartphone wands deployable as spinnable affordances. ACM SIG-MOBILE Mobile Computing and Communications Review, 16(4):4–5, Feb. 2013. DOI: 10.1145/2436196.2436199.

    We have built haptic interfaces featuring smartphones and tablets that use magnetometer-derived orientation sensing to modulate virtual displays, especially spatial sound, allowing, for instance, each side of a karaoke recording to be separately steered around a periphonic display. Embedding such devices into a spinnable affordance allows a “spinning plate”- style interface, a novel interaction technique. Either static (pointing) or dynamic (spinning) modes can be used to control “whirled” multimodal display, including a rotary motion platform, panoramic movies, and the positions of avatars in virtual environments.

  25. Julián Villegas and Michael Cohen. Roughness Minimization Through Automatic Intonation Adjustments. J. of New Music Research, 39(1):75–92, 2010. DOI: 10.1080/09298211003642480.

    We have created a reintonation system that minimizes measured roughness of parallel sonorities as they are produced. Intonation adjustments are performed by finding, within a user-defined vicinity, a combination of fundamental frequencies that yields minimal roughness. The vicinity imposition limits pitch drift and eases realtime computation. Prior knowledge of the temperament and notes being played is not necessary for the operation of the algorithm. We test a proof of concept prototype adjusting equal temperament intervals reproduced with a harmonic spectrum towards pure intervals in realtime. Pitch drift of the rendered music is not prevented but limited. This prototype exemplifies musical and perceptual characteristics of roughness minimization by adaptive techniques. We discuss the results obtained, limitations, possible improvements, and future work.

  26. Mohammad Sabbir Alam, Michael Cohen, Julián Villegas, and Ashir Ahmed. Narrowcasting for Articulated Privacy and Attention in SIP Audio Conferencing. J. of Mobile Multimedia, 5(1):12–28, 2009.

    In traditional conferencing systems, participants have little or no privacy, as their voices are by default shared with all others in a session. Such systems cannot offer participants the options of muting and deafening other members. The concept of narrowcasting can be applied to make these kinds of filters available in multimedia conferencing systems. Our system treats media sinks (in the simplest case, listeners) as full citizens, peers of the media sources (conversants’ voices), and we defined therefore duals of mute select: deafen attend, which respectively block a sink or focus on it to the exclusion of others. In this article, we describe our prototyped application, which uses existing standard Session Initiation Protocol (sip) methods to control fine-grained narrowcasting sessions. The runtime system considers the policy configured by the participants and provides a policy evaluation algorithm for media mixing and delivery. We have integrated a “virtual reality”-style interface with this sip backend to display and control articulated narrowcasting with figurative avatars.

Refereed conference articles

  1. Julian Dohmen and Julián Villegas. Hybrid approach to hrtf spatial upsampling. In Proc. 161st Audio Eng. Soc. Conv., Oct. 2026.

    We propose a hybrid method for spatial upsampling of head-related transfer function datasets. The method is based on the predictions of two known methods: the retrieval-augmented neural field (RANF) and the elevation-wise encoder-decoder (EWED), which have shown promising results in spatial upsampling datasets from only five measurements. The two methods have complementary error patterns: RANF’s retrieval-based prior gives it leading interaural time difference (ITD) accuracy but weaker log-spectral distortion (LSD), while EWED’s per-elevation autoencoders give good LSD results, but subpar ITD accuracy. We take the magnitude spectrum as a fixed convex combination of the predictions of the two methods and the ITD directly from RANF, reconstructing a minimum-phase impulse response. with the official LAP24 evaluation function. Our hybrid approach matches or reduces ITD error, ILD error, and LSD error relative to both baselines. upsampling methods can be successfully combined without retraining to improve spatial upsampling accuracy.

  2. Tom Krüger and Julián Villegas. Spline-based head-related transfer function compression. In Proc. 161st Audio Eng. Soc. Conv., Oct. 2026.

    Head-Related Transfer Functions (HRTFs) contain the spatial cues needed for immersive audio, but their size and the difficulty of personalizing them motivate compact, editable representations. We present a spline-based HRTF compression method that represents the magnitude spectrum by 43 control points per channel. One per Equivalent Rectangular Band (ERB) plus a high-frequency anchor, yielding a standardized format and achieving roughly 3.0:1 compression. Spectra are reconstructed with the Akima’s method, a parametric method, and with a hybrid method that selects the best spline per segment, as an upper bound. Across 302 subjects from the SONICOM database, is 3.94 dB and 3.55 dB for Akima’s and the hybrid method, respectively. A localization experiment with 18 participants, using a worst-case subject, found both reconstructions to be practically equivalent to the original HRTFs in polar angular error. In quadrant error, the equivalence held for Akima’s method but not for the hybrid. These results indicate that preserving the most prominent spectral feature per ERB is sufficient to retain cues needed for localization.

  3. Alaeddin Nassani, Yoshiko Ogawa, Ryuhei Yamada, Julián Villegas, Makiko Ohtake, Huidong Bai, and Mark Billinghurst. Toward VR teleoperation of lunar rovers: System design and field deployment at a lunar-analogue crater. In Proc. 5 Int. Conf. on Human-Computer Interaction in Space Exploration, Sep. 2026.

    We present a Virtual Reality (VR) teleoperation system that streams live 360∘ video and aligned RGB-Depth (RGB-D) point clouds from a tracked rover to a standalone Meta Quest 3 through an open-source split-rendering pipeline. We deployed and characterized the system outdoors at a lunar-analogue crater at the Fukushima Robot Testing Field: the 360∘ feed sustained its 30 fps source rate while RGB-D dropped about 2 fps; glass-to-glass latency was ∼0.6 s for 360∘ versus ∼2.1 s for RGB-D, requiring a 1.54 s synchronization delay that is an inherent cost of fusing the two streams. Outdoor sunlight and low-contrast analogue soil further degraded Time-of-Flight depth returns. Two small formative pilots (N = 3, N = 4) provided converging design signals: panoramic video was preferred over point-cloud and fused views, and a handheld controller over sustained hand-tracking gestures. These field measurements and indicative preferences establish a low-latency baseline and inform a planned confirmatory study under delayed-link conditions.

  4. Ikumi Saito and Julián Villegas. Comparing gammatone and mel-frequency cepstral coefficients for automatic breathy phonation classification. In Proc. of the Acoust. Soc. of Japan (ASJ) Autumn Meeting,, Sep. 2026.
  5. Li Tang and Julián Villegas. Japanese-accented English text-to-speech system. In Proc. IEEE Int. Conf. on Awareness Sci. and Tech (iCAST)., Sep. 2026.

    We present a Japanese-accented English text-to-speech (TTS) synthesizer based on fine-tuning a pretrained English TTS acoustic model with speech from Japanese learners of English. Objective evaluation based on predictions of Mean Opinion Scores (MOS) showed that the proposed TTS has a lower score than the original US English TTS (3.437 vs. 4.073). However, subjective evaluations based on transcription of sentences and accentedness rating, performed by native Japanese listeners, indicate that the proposed method yields lower Word Error Rate—WER (0.465 vs. 0.584) and higher Japanese-accentedness (71.2% vs. 20.7%). These results indicate that the fine-tuned system was perceived as more Japanese-accented and showed higher intelligibility, as measured by WER, among the 14 Japanese listeners in this experiment. We also propose alternatives for improving the speech quality of our method.

  6. Momoka Fujiwara, Tom Krüger, and Julián Villegas. Spatial sound rendering using spline-compressed head-related transfer functions. In Proc. Joint Conv. of Institutes of Electrical Engineer., Tohoku Section., Yamagata, Aug. 2026.

    Head-related transfer functions (HRTFs) are essential for 3D spatial audio, but datasets from multi-location measurements require enormous amounts of storage and computational resources. In this study, we present a system for efficient rendering of spatial audio from a dataset compressed using splines. This proposed system automatically extracts control points from a database, and uses Akima’s spline interpolation for reconstructing the magnitude spectrum of the HRTF at user-specified location, this can then be used to spatialize monaural audio. Our system achieves low-latency rendering while reducing data storage requirements and preserving audio quality.

  7. Ryoma Okuda and Julián Villegas. Zen-PCB: Material honesty and structural metaphor in a naked PCB granular looper instrument. In Proc. Int. Conf. New Interfaces for Musical Expression (NIME), London, UK, Jun. 2026.

    Zen-PCB is a musical instrument conceived as a counterpoint to the over-rationalized world and the passivity fostered by automated systems. It draws inspiration from ancient Japanese religious symbolism, the “kawaii” aesthetic, and the principle of material honesty: utilizing the naked printed circuit board as the instrument’s visual and tactile skin. Zen-PCB aims to reintroduce “ambiguity” and “spirituality,” qualities often absent in today’s functionalist designs. Zen-PCB encourages human initiative and celebrate the act of embracing “aimlessness.” It employs low-latency granular synthesis and destructive overdubbing. Zen-PCB features “matrix patching,” allowing users to interact with the PCB’s traces using conductive styluses. This interaction requires active participation, pushing back against the reliance on automated tools and sparking a creative tension between human and machine. Through structural metaphors, we reframe standard sampler functions as spiritual exercises. By incorporating the Buddhist concept of “impermanence” into its DSP architecture, Zen-PCB encourages users to engage with a continuous cycle of sonic creation and destruction. This pursuit of non-utilitarian experience offers a fresh perspective on creativity. We report on the design, implementation, and the reception of Zen-PCB, discussing how it transcends its function as a simple instrument, becoming a tool for physically embodying the cyclical nature of existence.

  8. Julián Villegas, Camilo Arévalo, Iain McGregor, Gerardo Sarria, Ethan Robson, and Juan Collazos-Mejía. Acoustic frequency and movement associations. In Joint conf. of 4th Iconicity Seminar (IcoSem) and 15th Int. Symp. on Iconicity in Language and Literature (ILL), Feb. 2026.

    We present a phonetically transcribed corpus of Air Traffic Control (ATC) communications from Japanese airports comprising about two hours of conversations between pilots and air traffic controllers. collected from LiveATC and Audio recordings were manually annotated by naive Japanese listeners with limited English proficiency (B2 or above), assisted with three offline Automatic Speech Recognition (ASR) systems: regular Whisper, Whisper fine-tuned with ATC transcriptions from mostly European airports, and Parakeet. For each annotator, a recording was presented along with a single ancillary transcription produced by one of these systems, so that all annotators were equally exposed to the three systems. Automatic and human annotations were compared with those made by an air traffic expert to establish Word Error Rates (WERs). Significant differences in WER were observed between the outputs of ASR systems, with fine-tuned Whisper achieving the best accuracy. Human transcriptions more than halved the WER of automatic ones, achieving similar performance among them regardless of the ASR system employed and English proficiency of the transcriber. Together, our findings suggest that ASR-assisted transcriptions could be effective in specialized jargon domains or when field experts are not available. In addition, the new publicly available dataset supports future research in aviation speech recognition.

  9. Yoshiki Sakai and Julián Villegas. Detection of Parkinsonian speech using auditory features. Proc. Mtgs. Acoust., 60(1):060016, 7 2026. DOI: 10.1121/2.0002345.

    This study aims to build a machine learning model to distinguish Parkinsonian speech (PS) from normal speech using models that reflect the transformations of sound throughout different stages in the auditory system. Parkinson’s disease is known to influence speech articulation for 89% of the patients. The impacted speech is often described as breathier and rougher compared to that of healthy individuals. Previous studies have predicted PS phonation using mel-frequency cepstral coefficients (MFCCs), Short-Time Fourier Transform and acoustic measures, such as F0, its jitter and shimmer.In this study, features related to the hearing process are explored, using outputs from auditory models. Since several phonation types can be perceived under the same category, despite production differences, our approach could potentially improve the accuracy of automatic PS classification. We compared predictions made with MFCCs, the output from a Gammatone filterbank, and from the Cascade of Asymmetric Resonators with Fast-Acting Compression (CARFAC) model (including its basilar membrane, inner hair cell potential, and neural activity patterns (NAPs) stages) using convolutional recurrent models (CNN-BiLSTM and CNN-BiGRU). Our results indicate that the CNN-BiGRU model with MFCCs achieves the best classification accuracy (72.06%), outperforming all auditory model-based features. NAPs yielded the best accuracy among auditory representations (71.17%).

  10. Li Tang, Veronica Khaustova, and Julián Villegas. Construction of a Japanese air traffic control communication corpus assisted with automatic speech recognition. In Proc. 159 Audio Eng. Soc. Conv., Oct. 2025.

    We present a phonetically transcribed corpus of Air Traffic Control (ATC) communications from Japanese airports comprising about two hours of conversations between pilots and air traffic controllers. collected from LiveATC and Audio recordings were manually annotated by naive Japanese listeners with limited English proficiency (B2 or above), assisted with three offline Automatic Speech Recognition (ASR) systems: regular Whisper, Whisper fine-tuned with ATC transcriptions from mostly European airports, and Parakeet. For each annotator, a recording was presented along with a single ancillary transcription produced by one of these systems, so that all annotators were equally exposed to the three systems. Automatic and human annotations were compared with those made by an air traffic expert to establish Word Error Rates (WERs). Significant differences in WER were observed between the outputs of ASR systems, with fine-tuned Whisper achieving the best accuracy. Human transcriptions more than halved the WER of automatic ones, achieving similar performance among them regardless of the ASR system employed and English proficiency of the transcriber. Together, our findings suggest that ASR-assisted transcriptions could be effective in specialized jargon domains or when field experts are not available. In addition, the new publicly available dataset supports future research in aviation speech recognition.

  11. Tom Krueger and Julián Villegas. Compression of head-related transfer functions using piecewise cubic Hermite interpolation. In Proc. of the 28th Int. Conf. on Digital Audio Effects (DAFx25), Ancona, Italy, Sep. 2025.

    We present a spline-based method for compressing and reconstructing Head-Related Transfer Functions (HRTFs) that preserves perceptual quality. Our approach focuses on the magnitude response and consists of four stages: (1) acquiring minimum-phase head-related impulse responses (HRIR), (2) transforming them into the frequency domain and applying adaptive Wiener filtering to preserve important spectral features, (3) extracting a minimal set of control points using derivative-based methods to identify local maxima and inflection points, and (4) reconstructing the HRTF using piecewise cubic Hermite interpolation (PCHIP) over the refined control points. Evaluation on 301 subjects demonstrates that our method achieves an average compression ratio of 4.7:1 with spectral distortion ≤ 1.0 dB in each Equivalent Rectangular Band (ERB). The method preserves binaural cues with a mean absolute interaural level difference (ILD) error of 0.10 dB. Our method achieves about three times the compression obtained with a PCA-based method.

  12. Ambrocio Gutierrez Lorenzo, Seunghun J. Lee, and Julián Villegas. Interaction of phonation and tone in San Miguel del Valle Zapotec. In Proc. 39 General Meeting of the Phonetic Society of Japan, Sep. 2025.
  13. Julián Villegas and Seunghun J. Lee. A meta-analysis of phonation research in the past 10 years. In Proc. 39 General Meeting of the Phonetic Society of Japan, Sep. 2025.
  14. Gabriele Ravizza, Julián Villegas, Christer P. Volk, Tore Stegenborg-Andersen, and Yan Pei. Target curve selection by differential evolution. In Proc. AES Int. Conf. on Headphone Technology, Espoo, Finland, Aug. 2025.

    We present an adaptive method to investigate preference within target frequency curve selection. An adaptive method, based on Differential Evolution, presents stimuli processed with frequency response curves from a population of candidate targets. The pool is iteratively updated based on each listener’s ratings in a paired comparison (with rating) paradigm. In the cycle of replacing the pool of candidates with more preferred curves, they converge towards each listener’s most preferred target curve. The final population of curves is then evaluated in a final multiple-stimulus comparison, and a ”winner” is found for each musical excerpt.

  15. Hui Shan Lai, Alaeddin Nassani, John Blake, and Julián Villegas. VR math bridge: Bridging interactivity in online education with AI and VR. In IEEE Gaming, Entertainment, and Media Conf., Taiwan, Jul. 2025.

    We present VR Math Bridge, a VR-based application designed to enhance calculus education by combining immersive virtual environments with AI-driven teaching assistance. VR Math Bridge creates a virtual classroom where students interact with Khan Academy videos and a 3D AI assistant that provides real-time, personalized feedback to their questions. This system leverages a floating panel for chapter selection, a virtual blackboard for video playback, and Cognitive 3D for analyzing user engagement. To demonstrate the system’s capabilities, we developed and deployed a prototype on Oculus Quest 3, focusing on derivatives as the initial test topic. Future directions include expanding to multi-user environments, enhancing visualization tools, and conducting comprehensive evaluations of its impact on learning outcomes.

  16. Alaeddin Nassani, Julián Villegas, and John Blake. Adaptive learning companions: Enhancing education with biosignal-driven digital human. In Proc. ACM Conf. on Human Factors in Computing Systems (CHI), Apr 2025. DOI: 10.1145/3706599.3719877.

    We introduce a novel approach to language learning leveraging digital humans as adaptive tutors within immersive XR environments. Our system’s novelty lies in the use of biosignals, specifically real-time heart rate data, collected from a Samsung Watch 7, to dynamically adapt the learning experience. The digital human tutor adjusts its behavior, feedback, and the difficulty of the learning content based on the learner’s inferred cognitive and emotional state. We present the fully developed system architecture, which integrates a customizable digital human powered by ConvAI, LLM, an XR environments, and a data streaming pipeline. While human participant testing is planned , preliminary insights from the system’s development demonstrate the technical feasibility of this approach. This research has the potential to significantly enhance language learning outcomes, engagement, and motivation by creating more personalized, and engaging learning experiences, paving the way for a new generation of adaptive educational technologies.

  17. Julián Villegas and Yusuke Sato. Perception of the missing fundamental in haptic complex tones. In Proc. 157 Audio Eng. Soc. Conv., Oct. 2024.

    We present a study on the perception of the missing fundamental in haptic complex tones. When asked to match an audible frequency to the frequency of a haptic tone with a missing fundamental frequency, participants in two experiments associated the audible frequency with lower frequencies than those present in the vibration, often corresponding to the missing fundamental of the haptic tone. This association was found regardless of whether the vibration was presented at the back (first experiment) or the feet (second experiment). One possible application of this finding could be the reinforcement of low frequencies via haptic motors, even when such motors may have high resonance frequencies.

  18. Yoshiki Sakai and Julián Villegas. Digital audio effect for simulating perceived sound of one’s voice. In Proc. Acoust. Soc. Japan, Autumn meeting, Osaka, Sep 2024.
  19. Julián Villegas, Ian Wilson, and Daichi Ishii. Inter-utterance rest position changes associated to Lombard speech. In Proc. Ultrafest XI, Jun. 2024.
  20. Patricia Fuente-García, Julián Villegas, and Irene de la Cruz Pavía. Only proficiency predicts older bilinguals’ word segmentation abilities in their second language. In Proc. 37 Annual Conf. on Human Sentence Process., Michigan, May 2024.

    Breaking down continuous speech into meaningful units is not a trivial task, yet we humans accomplish this feat effortlessly. A wealth of work has shown that adults and infants automatically integrate a number of cues available in the signal, and has revealed important changes in their relative weight and interplay across development (e.g. Mattys et al., 2005). However, a single study to date has examined whether their use remains constant throughout adulthood (Palmer et al., 2018). Similarly, the handful of studies available investigating the segmentation abilities of bilingual populations have exclusively focused on infants and young adults. To fill this gap in our knowledge, we examine older adults’ speech segmentation abilities, seeking to establish the impact of four central factors of bilingualism—age of acquisition, language proficiency, language use and frequency of language switch—in attaining native-like segmentation of their second language (L2). Specifically, we examine if and how these different bilingual dimensions determine their use of semantic, syntactic and phonological information in speech segmentation. A hundred and fifteen healthy older adults (+65y) from the Basque Country (Spain) participated in this study. All participants were native or native-like speakers of Spanish, and their knowledge of Basque ranged from minimal to native or native-like. All four bilingual dimensions were assessed using self-report questionnaires. Proficiency was additionally assessed in a grammar test and two standardized tasks (i.e. a lexical decision task and a naming task, de Bruin et al., 2017). Frequency of language switch was also assessed in a free switch task. In addition, a series of control tasks accounted for potential effects of cognitive health (MiniMental State Examination-37), fluid intelligence (Raven’s Progressive Matrices), and cognitive reserve (Cognitive Reserve Index Questionnaire). The segmentation task consisted of an adaptation to Basque of the one originally designed by Sanders and Neville (2000). Stimuli comprised 80 sentences, each presented in 3 versions (240 sentences in total): (1) “semantic” sentences were fully grammatical and meaningful, (2) “syntactic” sentences had all content words replaced with non-sense words while functors were preserved (as in Lewis Carroll’s Jabberwocky), and (3) “phonological” sentences exclusively contained non-sense words, only preserving phonological information (i.e. prosody, phonotactics and coarticulation). Each sentence was assigned a target phoneme that occurred only once, and the participants’ task consisted on indicating whether the target occurred word-initially or -medially. Multiple regression analysis demonstrated that proficiency (defined as the 1st component of a PCA including all measures) was the only bilingual factor which significantly predicted accuracy in the segmentation task (χ²(1)=67.08, p¡0.001). Moreover, they revealed significant interactions between proficiency and (1) sentence type (χ²(2)=373.40, p¡.001) and (2) target position (χ²(1)=3.87, p=.049). Post hoc comparisons revealed greater gains in accuracy as proficiency increased in semantic sentences, as compared with syntactic and phonological sentences (both p¡.001). In addition, they revealed larger proficiency related gains in accuracy in word-initial as compared with word-medial targets (p¡.001). These findings evidence a pivotal role of proficiency in word segmentation in older adult bilinguals and converge with previous literature showing different levels of plasticity across linguistic subsystems. The results also suggest that non-native speakers exhibit flexibility in using segmentation cues across the lifespan.

  21. Takuto Odaira and Julián Villegas. Delay digital audio effect inspired by acoustic environment on Mars. In Proc. AES 6th Int. Conf. on Audio for Games, Tokyo, Apr. 2024.

    Inspired by the recent discovery of two speeds of sound in the atmosphere of Mars, this article presents a real-time Digital Audio Effect (DAFX) implemented as a Virtual Studio Technology (VST) plugin. Our purpose is to simulate how a frequency-dependent speed transforms sound at different distances and to assess the degradation in speech intelligibility produced by such transformations. It was found that, at close distances, linear spectral changes could produce the largest intelligibility degradation, whereas, at long distances, degradation was similar to that introduced by other types of changes (logarithmic, sigmoid, or stepwise). The resulting plugin is not necessarily accurate, but can be used for expressive and illustrative purposes.

  22. Nicholas Carr, Julián Villegas, and John Blake. Dynamic online assessor: Scientific genre-specific automated feedback. In Proc. of JALTCALL Japan Association for Language Teaching Computer Assisted Language Learning Conf., 2024.

    The provision of written corrective feedback on language learners’ writing has been the focus of much investigation over the last three decades. Amongst this literature, there has been growing evidence that the provision of graduated feedback, which increases in explicitness as per learner needs in real time, can increase the learning potential of feedback on writing. However, the provision of such feedback has been limited to oral feedback due to its dialogical nature, with its implementation being reported as too time consuming for most real-life classrooms. This presentation introduces the ongoing development of DynaWrite (ver 1.0)—an online tool which automatically provides graduated feedback on learners’writing, thus overcoming the aforementioned modality and time restraints. Furthermore, in addition to detecting grammatical errors, the tool also provides automated dynamic feedback on brevity, clarity, objectivity and formality as they pertain to the genre of scientific writing. The proposed web application relays the input of a user to a chatbot based on a large language model. The chatbot is configured to increase the explicitness of feedback when the user is unable to resolve errors. We first briefly describe the theoretical foundation of the tool and graduated feedback. This is followed by a description of the four levels of feedback the tool provides, with level 1 being the most implicit and level 4 being the most explicit. We then discuss the process of error categorization and the process of instructing the chatbot to reliably detect these errors. The presentation concludes with a demonstration of how the tool is used and how it can track learner progress with specific error categories.

  23. Yoshiki Sato and Julián Villegas. Spectral tilt may have a smaller impact on the intelligibility of speech in noise. In Proc. IEEE Automatic Speech Recognition and Understanding Wkshp. (ASRU), Taipei, Taiwan, Dec. 2023. DOI: 10.1109/ASRU57964.2023.10389688.

    We compare spectral tilt modifications of plain speech made by Linear Predictive Coding—LPC transplantation and fractional roll-off filtering—FRF. These modifications were done in the scale of utterance-, segment-, or frame-based to equate the spectral tilt of speech produced in noise (Lombard speech) by the same speakers. In comparison with FRF, the LPC transplantation yielded larger objective speech intelligibility in noise gains when mixed with speech-shaped noise and larger spectral tilt errors as well. This finding suggests that the hitherto spectral tilt benefits assigned to Lombard speech are smaller than previously thought, especially for those reports based on segment- or frame-based LPC transplantation.

  24. Seunghun J. Lee, Julián Villegas, and Kunzang Namgyal. Tonal contrast in Drenjongke (Bhutia): an electroglottographic study. In Proc. 2nd Int. Conf. on Tone and Intonation, pages 19–23, Nov. 2023. DOI: 10.21437/TAI.2023-5.

    This paper investigates how phonation and tone interrelate in Drenjongke (Bhutia), a two-tone Tibeto-Burman language, by presenting results of electroglottography (EGG) data. EGG data of syllabary reading was collected from twelve Drenjongke speakers in Sikkim, India. After pruning the data, OQ measure- ments from 2,243 tokens were analyzed. The results were as follows: First, in vowel-only syllables, OQ values were higher in Low (L) tone than in High (H) tone. Second, in [+son] syl- lables, the sonorant portion shows higher OQ values in L-toned syllables than in H-toned ones; but the vowel portion does not. Third, after [-son] onsets, OQ values align with phonation; after aspirated onsets OQ values are higher than after modal onsets. Fourth, in voiced aspirates, which have tokens with negative or positive VOT, the regression line suggests that OQ values are higher when onsets have positive VOT in the first tercile. These four results suggest that phonation in Drenjongke has partial relevance to the realization tone depending on the nature of the tone: lexical or post-lexical.

  25. Edward Ly and Julián Villegas. Digital filter design via recurrent Cartesian genetic programming. In Proc. IEEE 13th Int. Wkshp. on Computational Intelligence & Applications (IEEE IWCIA2023), pages 7–12, Hiroshima, Japan, Nov. 2023. DOI: 10.1109/IWCIA59471.2023.10335891.

    We introduce a method for the automatic program induction of Single-Input/Single-Output (SISO) Infinite Impulse Response (IIR) filters for Digital Signal Processing (DSP) applications. Recurrent Cartesian Genetic Programming (RCGP) evolves a population of DSP programs, represented as directed cyclic/acyclic graphs, to generate a filter whose magnitude response approximates that of a target filter. The Log-Spectral Distance (LSD) is used as a fitness measure to minimize the differences between the magnitude responses of these filters, and the filter with the smallest distance is output. We evaluated our method by generating a number of filters, and found that the accuracy of the generated filters depends on both the order of the target filter and the parameter values set for the RCGP algorithm.

  26. Kevin Manuel Diaz España and Julián Villegas. Sonifying time series via music generated with machine learning. In Proc. 153 Audio Eng. Soc. Conv., Oct 2023.

    Conventional sonifications directly assign different aspects of data to auditory features and the results are not always “musical” as they do not adhere to a recognizable structure, genre, style, etc. Our system tackles this problem by learning orthogonal features in the latent space of a given musical corpus and using those features to create derivative compositions. We propose using a Singular Autoencoder (SAE) algorithm that identifies the most important Principal Components (PCs) in the latent space. As a proof-of-concept, we created sonifications of ionizing radiation measurements obtained from the Safecast project. Although the system successfully generates new compositions by manipulating the latent space, with each principal component changing different musical aspects, these changes may not be readily noticeable by listeners, despite the PCs being mathematically decorrelated. This finding suggests that higher-level features (such as associated emotion, etc.) may be needed for better results.

  27. Kazuki Fujita and Julián Villegas. Subjective evaluation of a virtual ad hoc loudspeaker array. In Proc. IEEE 12th Global Conf. on Consumer Electronics (IEEE GCCE), Nara, Oct 2023.

    We evaluated a virtual ad hoc audio spatializer. In this system, virtual loudspeakers and sound sources move arbitrarily on the horizontal plane in real-time. The spatializer, derived from the “a room within a room” spatialization method, comprises a listening area defined by the convex hull formed by the loudspeakers and a larger area (defined by a user) where audio sources can move. Compared to the direct spatialization of audio sources, subjective evaluations indicate that our method does not degrade the auditory location accuracy. However, virtual loudspeakers moving on a circular path yielded smaller ratings of changes in sound pressure level and location stability than when the loudspeaker motion was Brownian. For the latter, the stability rating of the virtual image is at a chance level. These results suggest that audio spatialization based on real ad hoc loudspeaker arrays could be achieved in the same manner but that the motion speed must be slow to avoid sound image instability.

  28. Julián Villegas, Kimi Akita, and Shigeto Kawahara. Psychoacoustic features explain subjective size and shape ratings of pseudo-words. In Proc. of Forum Acusticum, the 10 Conv. of the European Acoust. Assoc., Turin, Italy, Sep. 2023. DOI: 10.61782/fa.2023.0072.

    A previous study observed that the association between utterances and their ascribed meaning depends, among other factors, on phonation types. Specifically, the perceived size and shape of the referents of pseudo-words varied depending on whether they were pronounced with modal, creaky, falsetto or whisper phonation. In this study, we report a re-analysis of the same shape and size subjective ratings in terms of the four psychoacoustic measures: loudness, roughness, sharpness and pitch. We found that phonation types were positively associated with psychoacoustic features—falsetto with pitch, whisper with sharpness, and creakiness with loudness and roughness. Psychoacoustic features were also associated with subjective ratings: loudness and sharpness with subjective size and pitch and roughness with subjective shape. These findings indicate that psychoacoustic features (closer to the auditory representation of the heard sounds) may be good predictors of sound symbolism judgments.

  29. Camilo Arevalo and Julián Villegas. Study of auditory trajectories in virtual environments. In Proc. of Audio Mostly, Aug. 2023. DOI: 10.1145/3616195.3616210.

    A tool to study the apparent trajectories evoked by sounds (auditory trajectories) is presented. This tool is built with the aim of easing the task of the experimenters (building and analyzing interventions) and the task of the participants (reporting their opinions). By using infrared tracked controllers in a Virtual Reality environment, participants can freely describe the three-dimensional path evoked by a stimulus. The implemented tool also assists participants in recording trajectories by providing additional visual cues and feedback on the recorded data. A mock-up study is presented to demonstrate the benefits of the proposed system. Results from this study show that participants are able to accurately report elicited trajectories. While the implemented tool has limitations, such as the number of available blocks (only practice and main blocks), it could cover the needs of several laboratories. The tool is a valuable resource for researchers seeking to explore the perception and processing of auditory stimuli.

  30. Julián Villegas and Seunghun J. Lee. Creakiness judgments by Burmese and Vietnamese speakers. In Proc. 25 Conf. of the Oriental chapter of the Int. Committee for the Coordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA), Nov. 2022. DOI: 10.1109/O-COCOSDA202257103.2022.9997908.

    Vietnamese and Burmese listeners judged as creaky or not words in Du’an Zhuang, an unintelligible language for both groups. Although these languages differ in their tonal system, they use creakiness for phonemic contrast. The ratings of the two groups were similar, except in two non-creaky tones where the Vietnamese ratings were more likely to be non-creaky than the Burmese ratings. In terms of accuracy, the Vietnamese group was better doing these classifications. The differences between the two groups could indicate an influence of linguistic background in the determination of the creakiness of unintelligible words, however, it is also possible that the lack of control in the experimental setup (data was collected via an online survey) could have had a large effect on these results.

  31. Camilo Arevalo and Julián Villegas. Real-time spatialization via Eigen decomposition of head-related transfer functions. In Proc. 153 Audio Eng. Soc. Conv., New York, NY, Oct 2022.

    A binaural spatialization method based on a lossy compression of Head-Related Transfer function (HRTF) databases via Eigen decomposition is introduced. The high compression ratio achieved with this method allows for the inclusion of interpolated HRTFs (about a million) while reducing the resulting database 20% from the original size. Compressed HRTFs have a spectral distortion ≤ 1 dB between 0.1 and 16 kHz. This distortion seems to have no significant effect on subjective accuracy, which is similar to the accuracy achieved with currently used audio spatializers. Compared to one of these methods, the proposed method shows an improvement in terms of processing time and memory requirements.

  32. Edward Ly and Julián Villegas. Additive synthesis via recurrent Cartesian genetic programming in FAUST. In Proc. 153 Audio Eng. Soc. Conv., New York, NY, Oct 2022.

    We introduce a novel method for generating arbitrary additive synthesizers for audio replication: given a target audio signal, of input-output signals), digital signal processes, expressed as cyclic graph structures, are evolved via Recurrent Cartesian Genetic Programming (RCGP) to generate an approximation of the target audio arbitrary accuracy. The Log-Spectral Distance (LSD) between the target and actual output signals is used to measure the fitness of each solution candidate. The candidate with the smallest LSD is returned as the solution. A working implementation of RCGP in C that generates Functional AUdio STream (FAUST) code is provided, then evaluated via re-synthesis of steady state tones as provided in the Sandell Harmonic ARChive (SHARC). Early results show promise as a mean LSD value of 5.1 dB was achieved after 1000 generations, with 0.52 dB being the minimum LSD found.

  33. Camilo Arévalo and Julián Villegas. Compressing head-related transfer function databases by Eigen decomposition. In IEEE Int. Wkshp. on Multimedia Signal Processing (MMSP), Tampere, Finland, Sep 2020. DOI 10.1109/MMSP48831.2020.9287134.

    A method to reduce the memory footprint of Head-Related Transfer Functions (HRTFs) is introduced. Based on an Eigen decomposition of HRTFs, the proposed method is capable of reducing an a database comprising 6,344 measurements from 36.30 MB to 2.41 MB (about a 15:1 compression ratio). Synthetic HRTFs in the compressed database were set to have less than 1 dB spectral distortion between 0.1 and 16 kHz. The differences between the compressed measurements with those in the original database do not seem to translate into degradation of perceptual location accuracy. The high degree of compression obtained with this method allows the inclusion of interpolated HRTFs in databases for easing the real-time audio spatialization in Virtual Reality (VR).

  34. Julián Villegas. Spatial perception of Risset notches. In Proc. 14 Int. Symp. on Computer Music Multidisciplinary Research (CMMR), Marseille, Oct. 2019. HAL hal-02382500.

    The apparent movement of Risset tones and Risset notches (i.e., the opposite of Risset tones, replacing silence by noise and frequency components by notches) are investigated. Contrary to previous findings, no significant differences between horizontal and vertical movement associations were found, regardless of stimuli. However, whereas the tones were more likely to be subjectively associated with approaching sources, notches were associated with receding ones. The direction of frequency glide also had a significant effect: ascending glides were more likely to be associated with horizontal movements and descending ones with vertical movements. These findings suggest that although both stimuli evoke similar illusions, the perception of their spatial attributes are different.

  35. Julián Villegas and Seunghun J. Lee. Creating maps for linguistic field-work using R. In Proc. 6 Int. Conf. on Language Documentation & Conservation (ICLDC), Feb 2019.

    In this showcase we will introduce a step-by-step method to create maps in R to be used as reference in field-work publications. At the end of the session, attendees should be able to produce maps highlighting administrative districts and list of cities relevant to their work, as well as insets to show the context of the region of interest.

  36. Jeremy Perkins, Seunghun J. Lee, Julián Villegas, and Kosei Otsuka. Using psychoacoustic roughness to measure creakiness in Burmese. In Proc. 5 NINJAL Int. Conf. on Phonetics and Phonology, Oct 2018.

    In this paper we found that psychoacoustic roughness and not spectral tilt reliably distinguished creaky from non-creaky tones in Burmese. These results match Gruber’s (2011) findings: 1. Creakiness is found late in syllables in checked and creaky tones, 2. Creakiness is only found in words read in isolation, and not in frame sentences.

  37. Julián Villegas. Association of frequency changes with perceived horizontal and vertical movement. In Proc. Int. Symp. on Universal Acoustical Communication, Oct. 2018.

    Frequency changes are often associated with changes in the apparent height or elevation of a sound source (a phenomenon referred as the Pratt effect). Situations arise where these frequency changes are associated with approaching sources from the horizon instead (the so called, Doppler illusion—not to be confused with the Doppler effect). This possible contradiction was investigated with a series of experiments where subjects were asked to state the apparent origin (above or at the horizon) and movement (approaching or receding) of Risset tones (ascending or descending glides of 20 frequency components separated half an octave) presented diotically in a forced choice manner. Results of such experiments indicate that when giving subjects an alternative to choose apparent origin and movement, downward Risset tones are most likely to be associated with sources receding horizontally and upward Risset tones with sources approaching also horizontally. Associations with vertical movements (approaching or receding from above) were less likely. These findings suggest that the Doppler illusion is stronger than the Pratt effect, at least for Risset tones.

  38. Camilo Arevalo, Gerardo M. Sarria M., and Julián Villegas. Accurate spatialization of VR sound sources in the near field. In Audio Engineering Society Int. Conf. on Spatial Reproduction - Aesthetics and Science, Jul 2018.

    The aim of this research is to offer a spatialization alternative in VR engines (such as Unity) that bridges the gap left by regular spatializers for sounds in the near-field. In this paper, we detail the development of HRIR_U, an API for spatialization of multiple sources in the near-field in real-time, and present a comparison between perceived distance accuracy achieved with native methods in Unity and that achieved with HRIR_U. This comparison indicates that HRIR_U could offer better spatialization, though scalability and performance should be improved.

  39. Anh T. Pham, Truong Cong Thang, Julián Villegas, and Michael Cohen. VLC-based smart supermarket (SMARTKet): Key concepts and enabling technologies. In Proc. IEEE 6th Global Conf. on Consumer Electronics (GCCE), Nagoya, Japan, 2017. DOI 10.1109/GCCE.2017.8229248.

    We present the key concepts and design of our proposed framework for a smart supermarket (SMARTKet). We briefly introduce the infrastructure, smart functions, and enabling technologies of the SMARTKet implementation. We especially focus on the basic principles, performance evaluation in terms of localization accuracy, and proof-of-concept implementation of the indoor navigation system using visible light communications (VLC) localization technology in the context of SMARTKet.

  40. Julián Villegas and Shoma Saito. Assisting system for grocery shopping navigation and product recommendation. In Proc. IEEE 6th Global Conf. on Consumer Electronics (GCCE), Nagoya, Japan, 2017. DOI 10.1109/GCCE.2017.8229387.

    We present the key concepts and design of our proposed framework for a smart supermarket (SMARTKet). We briefly introduce the infrastructure, smart functions, and enabling technologies of the SMARTKet implementation. We especially focus on the basic principles, performance evaluation in terms of localization accuracy, and proof-of-concept implementation of the indoor navigation system using visible light communications (VLC) localization technology in the context of SMARTKet.

  41. Julián Villegas and Takaya Ninagawa. Pure-data-based transaural filter with range control. In Proc. 5th Int. Pure Data Convention, Nov 2016.

    We present an extension to Pure-data by which users can truly spatialize sound via a pair of loudspeakers, i.e., spatialize monaural sound sources at an arbitrary azimuth, elevation, and distance. Although transaural techniques have been long explored, our system takes advantage of a recently collected Head-Related Impulse Response (hrir) dataset measured in the near field (20–160 cm from the center of a mannequin’s head) to allow a more accurate distance control, a missing feature in other implementations.

  42. Julián Villegas, Tore Stegenborg-Andersen, Nick Zacharov, and Jesper Ramsgaard. Effect of presentation method modifications on standardized listening tests. In Proc. 141 Audio Eng. Soc. Int. Conv., Sep. 2016.

    This study investigates the impact of relaxing presentation methods on listening tests by comparing results from two identical listening experiments carried out on two countries and comprising two presentation methods: the ITU-T P.800 Absolute Category Rating (ACR) recommendation and a modified version of it where assessors had more control on the reproduction of the samples. Compared with the standard method, test duration was reduced on average 37used on the ratings of codecs were found, but a significant effect of site on ratings and duration were found. We hypothesize that in the latter case, cultural differences and instructions to the assessors could explain these effects.

  43. Jeremy Perkins, Seunghun Lee, and Julián Villegas. The roles of phonation and f0 in Wuming Zhuang tone. In Proc. 22nd Himalayan Languages symp., Jun. 2016.

    This study reports phonetic measurements of the tonal system of Wuming Zhuang. While previous analyses have described Wuming Zhuang’s tone contrasts using F0 only, this study finds that that creaky phonation can distinguish pairs of tones that have similar F0 contours, suggesting that creakiness, in addition to F0, may play a role in distinguishing tones. A composite acoustic algorithm is applied as a way to compute creaky phonation and is offered as an alternative method for linguists interested in measuring phonation from the acoustic signal.

  44. Donna Erickson, Julián Villegas, Ian Wilson, Yuki Iguro, Jeff Moore, and Daniel Erker. Some acoustic and articulatory correlates of phrasal stress in Spanish. In Proc. 8 Speech Prosody, Boston, MA, May 2016. DOI 10.21437/SpeechProsody.2016-92.

    All spoken languages show rhythmic patterns. Recent work with a number of different languages (English, Japanese, Mandarin Chinese, French) suggest that metrically assigned stress levels of the utterance show strong correlations with the amount of jaw displacement, and corresponding F1 values. This paper examines some articulatory and acoustic correlates of Spanish rhythm; specifically, we ask if there is a correlation between phrasal stress values metrically assigned to each syllable with acoustic and articulatory values. We used video recordings of 3 Salvadoran Spanish speakers to measure for each vowel maximum jaw displacement, mean F0, mean intensity, mean duration, and mid vowel F1 of two Spanish sentences. The results show weak but significant correlations between jaw displacement and F1/ intensity, but no correlation between jaw displacement and F0. We also found strong correlations between stress, duration, and F1, and weaker, but significant correlations between stress and mean intensity /maximum jaw displacement.

  45. Jeremy Perkins, Seunghun Lee, and Julián Villegas. An interplay between F0 and phonation in Du’an Zhuang tone. In Proc. 5 Int. Symp. on Tonal Aspects of Languages, pages 56–59, Buffalo, May 2016. DOI: 10.21437/TAL.2016-12.

    This paper undertook an acoustic study of the tone system of Du’an Zhuang, finding that unlike the standard dialect, Wuming Zhuang, its tone system involved phonation differences in addition to F0 and duration differences. It was found that two of the six tones in unchecked syllables in Du’an Zhuang involved significant creakiness near the midpoint of the vowel. In checked syllables, a three-way tonal contrast was observed based on F0 contours, but not creakiness. These results suggest a phonological tone contrast that involves both F0 and creakiness. Among pairs of tones that differed in their phonation, significant differences in the timing of F0 fall were discovered. Additionally, the two creaky tones differed in the timing of the maximum creakiness. Future research on the perception side could establish whether and to what extent Du’an Zhuang speakers utilize creakiness and F0, and their relative timing, in discerning between tonal categories.

  46. Julián Villegas. An online benchmarking platform for visualizing ionizing radiation doses in different cities. In Proc. of IEEE EATIS: 8th Euro-American Conf. on Telematics and Information Systems, April 2016. DOI 10.1109/EATIS.2016.7520144.

    A working prototype for alternative visualizations of environmental data (currently, ionizing radiation) measured with bGeigie nano Safecast sensors is presented. Contrary to previous interfaces, in this visualization users have finer control of the displayed data (i.e., can determine date ranges, compare locations, decide the averaging areas, etc.) and more detailed information of the resulting visualization (size of the samples per day and per region, etc.). With this new data visualization, it is easier to compare local environment figures with those of other regions of the planet.

  47. Taku Nagasaka, Shunsuke Nogami, Julián Villegas, and Jie Huang. Influence of spectral energy distribution on elevation judgments. In Proc. 139 Audio Eng. Soc. Int. Conv., Oct. 2015.

    The relative influence of spectral cues on elevation localization was investigated by comparing judgements of loudspeaker reproduced stimuli spatialized with three methods: 3D vector-based amplitude panning (3D- vbap), 2D-vbap in conjunction with hrir convolution, and equalizing the stimuli to simulate spectral peaks and notches naturally occurring at different angles (equalizing filters). For the last two methods a single horizontal loudspeaker array was used. As expected, smallest absolute errors were observed in the vbap judgements regardless of presentation azimuth; no significant difference in the mean absolute error was found between the other two methods. But, for most presentation azimuths, the method based on equalizing filters yielded less dispersed results. These results could be used for improving elevation localization in two-dimensional vbap reproduction systems.

  48. Shunsuke Nogami, Taku Nagasaka, Julián Villegas, and Jie Huang. Influence of spectral energy distribution on subjective azimuth judgements. In Proc. 139 Audio Eng. Soc. Int. Conv., Oct. 2015.

    In this research, we compare subjective judgements of azimuth obtained by three methods: Vector-Based Amplitude Panning (vbap), vbap mixed with binaural rendition over loudspeakers (vbap+hrtf), and a newly proposed method based on equalizing spectral energy. In our results, significantly smaller errors were found for the stimuli treated with vbap and hrtfs; differences between the other two treatments were not significant. Regarding spherical dispersion of the judgements, vbap results have the greatest dispersion, whereas the dispersion on the results of the other two methods were significantly smaller, however similar between them. These results suggest that horizontal localization using vbap methods can be improved by applying a frequency dependent panning factor a opposed to a constant scalar as commonly used.

  49. Shunsuke Nogami, Taku Nagasaka, Julián Villegas, and Jie Huang. Improvement of azimuth perception in single-layer speaker array systems. In Proc. Acoust. Soc. Japan, Autumn meeting, Aizu Wakamatsu, Japan, Sep. 2015. In Japanese.
  50. Tomomi Sugasawa, Jie Huang, and Julián Villegas. Relative influence of spectral bands in horizontal-front localization of white noise. In Proc. 137 Audio Eng. Soc. Conv., Oct. 2014.

    The relationship between horizontal-front localization and energy in different spectral bands is investigated in this research. Specifically, we tried to identify which spectral regions produced changes in the judgments of the position of a white noise when each band was removed from a front loudspeaker and presented via side loudspeakers. These loudspeakers were set at left and right from the front-midsagittal plane of the listener. Participants were asked to assess whether the noise was coming from the front loudspeaker as bands were moved from front to side loudspeakers. Results from a pilot study suggested differences in the relative importance of spectral bands for horizontal-front localization.

  51. Julián Villegas. Movement perception of Risset tones with and without artificial spatialization. In Proc. 137 Audio Eng. Soc. Conv., Oct. 2014.

    The apparent radial movement (approaching or receding) of Risset tones was studied for sources in front, above, and to the right of listeners. Besides regular Risset tones, two kinds of spatialization were included: global (regarding the tone as a whole) and individual (spatializing each of its spectral components). The results suggest that regardless of the direction of the glissando, subjects tend to judge them as approaching. The effect of spatialization type was complex: For upward Risset tones, judgements were, in general, aligned with the direction of the spatialization, but this was not observed in the downward Risset tones. Furthermore, individual spatialization yielded judgements comparable to those of non-spatialized stimuli, whereas spatializing the stimuli as a whole yielded judgments more aligned with the treatment.

  52. Wataru Sanuki, Julián Villegas, and Michael Cohen. Spatial sound for mobile navigation systems. In Proc. 136 Audio Eng. Soc. Conv., 2014.
  53. Michael Cohen, Rasika Ranaweera, Kensuke Nishimura, Yuya Sasamoto, Shun Endo, Tomohiro Oyama, Tetunobu Ohashi, Yukihiro Nishikawa, Ryo Kanno, Anzu Nakada, Julián Villegas, Yong Ping Chen, Sascha Holesch, Jun Yamadera, Hayato Ito, Yasuhiko Saito, and Akira Sasaki. “Tworlds”: Twirled worlds for multimodal ‘padiddle’ spinning & tethered ‘poi’ whirling. In Proc. of SIGGRAPH, pages 67:1–67:1, Nov. 2013.

    Modern smartphones and tablets have magnetometers that can be used to detect yaw, which data can be distributed to adjust ambient media. Either static (pointing) or dynamic (twirling) modes can be used to modulate multimodal displays, including 360∘ imagery and virtual environments. Azimuthal tracking especially allows control of horizontal planar displays, including panoramic and turnoramic imaged-based rendering, spatial sound, and the position of avatars, virtual cameras, and other objects in virtual environments such as Alice, as well as rhythmic renderings such as musical sequencing.

  54. Michael Cohen, Rasika Ranaweera, Kensuke Nishimura, Yuya Sasamoto, Tomohiro Oyama, Tetsunobu Ohashi, Anzu Nakada, Julián Villegas, Yong Ping Chen, Sascha Holesch, Jun Yamadera, Hayato Ito, Yasuhiko Saito, and Akira Sasaki. Twirled affordances, self-conscious avatars, & inspection gestures. In Proc. SIGGRAPH Asia: Symposium on Mobile Graphics and Interactive Applications, pages 95:1–95:1, Nov. 2013.

    Contemporary smartphones and tablets have magnetometers that can be used to detect yaw, which data can be distributed to adjust ambient media. We have built haptic interfaces featuring smartphones and tablets that use compass-derived orientation sensing to modulate virtual displays. Embedding mobile devices into pointing, swinging, and flailing affordances allows “padiddle”-style interfaces, finger spinning, and “poi”-style interfaces, whirling tethered devices, for novel interaction techniques.

  55. Yuya Sasamoto, Michael Cohen, and Julián Villegas. Controlling spatial sound with table-top interface. In Proc. IEEE Int. Joint Conf. on Awareness Science and Technology & Ubi-Media Computing, pages 713–718, Nov. 2013.

    Interactive table-top interfaces are multimedia devices which allow sharing information visually and aurally among several users. Table-top interfaces for spatial sound environments are frequently investigated in the field of the human interfaces. Table-top interfaces are utilized as groupware and it is suitable for collaborative work, and it is convenient for a group working on theme related to sound systems. A representative of table-top musical instrument is the reacTable. In this paper, we present a way to control the position of multiple sounds in a spatial sound environment via a table-top interface. Sound localization is required to discriminate and recognize clearly sounds. We have been investigating musical table-top instruments which are capable of controlling multiple sound in spatial sound environments. One of the main features of this new developed system is that multiple users can control the spatialization of independently sounds in real-time. We verified changes of user recognition to multi-sound with a spatial sound environment.

  56. Julián Villegas and Michael Cohen. Real-time head-related impulse response filtering with distance control. In Proc. 135 Audio Eng. Soc. Conv., Oct. 2013.

    We present a new software application based on a recently collected hrir database comprising measurements at different distances. The new application, programmed in Pure-data, is capable of directionalizing sound objects at any azimuth, at elevations between -40 degrees and 90 degrees, and at distances 20–160 cm. This truly 3D spatialization is done by pre-calculating the minimum-phase version of the HRIRs and computing the interpolation of a maximum of four hrir measurements, depending upon the virtual location. In the same way, interaural time differences are computed and applied to the convolved signal. For demanding real-time constraints, the number of taps used for the convolution can be adjusted, up to a maximum of 1024.

  57. Jorge Gonzalez Alonso and Julián Villegas. Dominance takes precedence: L3 English processing by Basque-Spanish bilinguals. In AESLA, 31 Int. Conf. on Communication, Cognition and Cybernetics, page NA, 2013.

    Word-formation processes vary greatly among languages, although those which are typologically close tend to cluster around particular configurations which may or may not differ from those of other linguistic families. Compound words in Romance and Germanic languages have been considered by both theoretical linguists and acquisitionists, with the latter focusing more on the interplay between two or more systems in a multilingual setting. The case of deverbal N+N compounds (e.g. can opener) in English as compared to their [V+N]N Spanish semantic equivalents (e.g. abrelatas ‘can opener’, lit. ‘opens-cans’) is particularly interesting. What seems apparent is that Spanish and English do not lexicalise verb-noun relationships in the same way. Basque, in contrast, does seem to have direct parallels with English: Basque deverbal compounds are also right-headed N+N constructions, in which the deverbal head has been nominalised through affixation (e.g. lata irekigailu, lit. ‘can opener’). Considering these facts, are there any facilitatory effects in processing for those bilinguals whose L1 is similar to the L3 (English) in the formation of deverbal compounds? An experiment was carried out in which we controlled for both language profile and proficiency. We predicted practically equal accuracy rates for all groups at comparable levels of proficiency, since the effect is not expected to override lexical knowledge; a faster performance of the monolingual group, due to an attested higher processing cost in bilinguals (Ivanova & Costa, 2008); and shorter response latencies for the Basque-dominant bilinguals as opposed to their Spanish-dominant counterparts, since the critical structure is hypothesised to be more readily available for the former group. Response latencies and accuracy rates were analysed with two independent two-way ANOVA with proficiency in English and language profile as factors. Results have largely matched our predictions.

  58. Michael Cohen, Rasika Ranaweera, Kensuke Nishimura, Yuya Sasamoto, Yukihiro Nishikawa, Tetunobu Ohashi, Ryo Kanno, Tomohiro Oyama, Anzu Nakada, and Julián Villegas. Whirled Sequencing of Spatial Music. In Proc. Audio Eng. Soc. Japan Sect. Conf., Sendai, Oct. 2012.

    “Poi,” originally a Maori performance art featuring whirled tethered weights, combines elements of dance and juggling. It has been embraced by contemporary festival culture (especially rave-style electronic music events), including extension to “glowstringing,” in which a glow stick (chemiluminescent plastic tube) is whirled at the end of a string. We further modernize this activity, opening it up to internet-amplified multimedia. The ubiquity of the contemporary smartphone makes it an attractive platform for even location-based attractions. By sensing its magnetometer, the twirling of a mobile phone can be used to sequence score-following music. Synchronizing this sequencing with sound spatialization, also modulated by the azimuth of the whirled phone, as through an annular (ring-shaped) speaker array, allows interactive, multimodal interaction.

  59. Elizabeth Godoy, Yannis Stylianou, and Julián Villegas. Unsupervised Normal-to-Lombard Spectral Envelope Transformation; Examining Loudness, Voicing & Stationarity. In Proc. The Listening Talker Wkshp., Edinburgh, May 2012.

    When speaking in noisy environments, humans modify their speech in order to make it more intelligible: this phenomenon is known as the Lombard effect [1],[2]. It has been shown that, among the various Lombard modifications, those to the spectral envelope account for the largest increases in speech intelligibility [3]. The present work examines and seeks to exploit the spectral envelope differences between Normal and Lombard speech for multiple (4 male, 4 female) speakers of the GRID corpus in an unsupervised context, i.e., in the absence of segmentation or phonetic labeling. Our goals are twofold: 1) to transform the Normal speech spectral envelope towards that of the Lombard; 2) to isolate acoustic criteria that help to identify and better understand important (e.g. perceived) spectral differences between Normal and Lombard speech.

  60. Julián Villegas, Martin Cooke, and Catherine Mayo. The role of durational changes in the Lombard speech advantage. In Proc. The Listening Talker Wkshp., Edinburgh, UK, May 2012.

    Speech produced in the presence of noise – Lombard speech (LS) – has been found to be more intelligible than ’normal’ speech when presented in equivalent amounts of noise. However, the origin of the Lombard speech advantage remains unclear. Part of the benefit appears to stem from spectral changes in LS which shift energy into the 1-4 kHz region where it better escapes energetic masking by speech-shaped noise. Other parameters which show changes in LS include F0 and duration. Lu & Cooke (2009) modified the mean F0 and spectrum (both independently and jointly) of normal speech, demonstrating a clear advantage of spectral modification but no effect of F0. The current study extended Lu & Cooke (2009) in two directions. First, durational modifications to reflect differences between normal and Lombard speech were included. Second, as well as global (per utterance) changes, local (per frame) modifications were applied. Four male and four female talkers produced simple sentences containing spoken letter and number keywords in quiet and in intense speech-shaped noise (96 dB SPL). A perception experiment explored global versus local modifications of spectral and durational parameters applied independently. Spectral changes based on global or local spectral difference measures were equally beneficial. Taken across all talkers, durational changes to normal speech produced no intelligibility benefit, whether applied globally (i.e. linear stretching or compression) or locally (using dynamic time warping to align normal and Lombard frames). However, when talkers were partitioned in two groups according to their speech rate in noise, some effect of durational modifications was observed: normal speech modified to the faster speech rate group was significantly less intelligible than unmodified speech, while conversely for the group with slower Lombard speech a small intelligibility benefit was present. These findings suggest that durational differences between normal and Lombard speech can affect intelligibility, but the benefit or otherwise depends on individual differences in speech rate.

  61. Julián Villegas and Martin Cooke. Maximising objective speech intelligibility by local f0 modulation. In Proc. Interspeech, pages 1704–1707, 2012. DOI: 10.21437/Interspeech.2012-466.

    We investigated the effect on objective speech intelligibility of scaling the fundamental frequency (f0) of voiced regions in a set of utterances. The frequency scaling was driven by maximising the glimpse proportion in voiced epochs, inspired by musical consonance maximisation techniques. Results show that depending on the energetic masker and the signal to noise ratio, f0 modifications increased the mean glimpse proportion by up to 15%. On average, lower mean f0 changes resulted in greater glimpse proportions. It was also found that the glimpse proportion could be a good predictor of music consonance.

  62. Vincent Aubanel, Julián Villegas, and Martin Cooke. Conversing in the presence of another conversation: interactive and Lombard effects. In Wkshp. on the production and comprehension of conversational speech, Nijmegen, the Netherlands, Dec. 2011.

    Conversational speech is usually characterized by well described departures in production from carefully pronounced speech, while it maintains a fair level of comprehension. Less is known however on the speech production modifications induced by a background conversation on a foreground conversation, and how speakers manage to maintain intelligibility and comprehension in this scenario. Extending a previous study with Spanish native speakers, we recorded pairs of English native speakers engaging in natural dialogues in the absence or presence of another talker pair. We observed small but significant increases in energy and F1 across conversations, and larger prosodic effects (increase of f0, decrease in speech rate) within dialogues, which are not easily explained in purely energetic masking terms. Indeed, background conversations are different from noise used in traditional Lombard studies in that they consist of intelligible speech, which creates the potential for informational masking at the ears of the interlocutor. We further tested whether speakers’ eye-contact with their interlocutor could influence their capacity to cope for the disruptive effect of a background conversation. Preliminary result point to an attenuation of Lombard effects (i.e., intensity, f0, F1, speech rate) in the absence of eye-contact condition, which could be the sign of an active monitoring of the informational content of the conversation, thus further minimizing the masking effect of the background conversation.

  63. Michael Cohen, Rasika Ranaweera, Hayato Ito, Shun Endo, Sascha Holesch, and Julián Villegas. Whirling interfaces: Smartphones & tablets as spinnable affordances. In Proc. of ICAT—The 21 Int. Conf. on Artificial Reality and Telexistence, Osaka, Nov. 2011.

    Interfaces featuring smartphones and tablets that use magnetometer-derived orientation sensing can be used to modulate virtual displays. Embedding such devices into a spinnable affordance allows a “spinning plate”-style interface, a novel interaction technique. Either static (pointing) or dynamic (whirled) mode can be used to control multimodal display, including panoramic and turnoramic images, the positions of avatars in virtual environments, and spatial sound. “Spinning,” in which a flatish object is whirled with an extended finger or stick is a disappearing art. We hope to re-motivate this vanishing skill, modernizing it and opening it up to internet-amplified multimedia. The ubiquity of the modern smartphone makes it an attractive platform for even location-based attractions. We are experimenting with embedding mobile devices into suitable affordances that encourage their spinning. Using azimuthal (yaw) tracking especially allows such devices to control horizontal planar displays such as periphonic spatial sound, as well as avatar heading and (QTVR-style) panoramic and turnoramic imaged-based rendering.

  64. Michael Cohen and Julián Villegas. From Whereware to Whence- and Whitherware: Augmented Audio Reality for Position-Aware Services. In Proc. Int. IEEE Symp. on VR Innovation (ISVRI), Singapore, March 2011.

    Since audition is omnidirectional, it is especially receptive to orientation modulation. Position can be defined as the combination of location and orientation information. Location-based or location-aware services do not generally require orientation information, but position-based services are explicitly parameterized by angular bearing as well as place. “Whereware” suggests using hyperlocal georeferences to allow applications location-awareness; “whence- and whitherware” suggests the potential of position-awareness to enhance navigation and situation awareness, especially in realtime high-definition communication interfaces, such as spatial sound augmented reality applications. Combining literal direction effects and metaphorical (remapped) distance effects in whence- and whitherware position-aware applications invites over-saturation of interface channels, encouraging interface strategies such as audio windowing, narrowcasting, and multipresence.

  65. Vincent Aubanel, Martin Cooke, Julián Villegas, and Maria Luisa Garcia Lecumberri. Conversing in the presence of a competing conversation: effects on speech production. In Proc. Interspeech, 2011.

    This study investigates how a background conversations affect foreground conversations, and how speakers may adjust their speech to overcome the perturbations. Three pairs of speakers were recorded in different combinations of simultaneous dialogues, and speech production modifications were investigated at an acoustical and interactional level. In addition to displaying standard Lombard effects, speakers were found to produce less back-channels and more interruptions in the presence of a background conversation. A decrease in the precision of turn taking was also observed. These results provide a better understanding of the strategies speakers may be developing in dealing with a concurrent conversation in the view of incorporating them into spoken dialogue systems.

  66. Julián Villegas, Martin Cooke, Vincent Aubanel, and Marco A. Piccolino-Boniforti. MTRANS: A multi-channel, multi-tier speech annotation tool. In Proc. Interspeech, 2011.

    MTRANS, a freely available tool for annotating multi-channel speech is presented. This software tool is designed to provide visual and aural display flexibility required for transcribing multi-party conversations; in particular, it eases the analysis of speech overlaps by overlaying waveforms and spectrograms (with controllable transparency), and the mapping from media channels to annotation tiers by allowing arbitrary associations between them. MTRANS supports interoperability with other tools via the Open Sound Control protocol.

  67. Julián Villegas and Michael Cohen. Hrir~: Modulating range in headphone-reproduced spatial audio. In Proc. of 9th ACM SIGGRAPH Int. Conf. on VR Continuum and Its Applications in Industry, pages 89–84, Dec. 12–13 2010. DOI: 10.1145/1900179.1900198.

    Hrir~, a new software audio filter for Head-Related Impulse Response (hrir) convolution is presented. The filter, implemented as a Pure-Data object, allows dynamic modification of a sound source apparent location by modulating its virtual azimuth, elevation, and range in realtime. The last attribute being missing in surveyed similar applications. With hrir~ users can virtually localize monophonic sources around a listener’s head in a region delimited by elevations between [−40,90]°, and ranges between [20,160] cm from the center of the virtual listener’s head. An application based on hrir~ is presented to illustrate its benefits.

  68. Julián Villegas and Michael Cohen. “Gabriel”: Geo-Aware Broadcasting for In-vehicle Entertainment and Larger Safety. In Proc. 135 Audio Eng. Soc. Int. Conv., October 2010.

    We have retrofitted a vehicle with location-aware advisories/announcements, delivered via wireless headphones for passengers and bone-conduction headphones for the driver. Our prototype differs from other research in the spatialization of the aural information. Besides the commonly used landmarks to trigger audio streams delivery, our prototype uses geo-located virtual sources to synthesize the spatial soundscapes. Intended as a “proof of concept” and testbed for future research, our development features multilingual tourist information, navigation instructions, and traffic advisories rendered simultaneously.

  69. Julián Villegas, Michael Cohen, Ian Wilson, and William Martens. Influence of Roughness on Preference of Musical Intonation. In Proc. 128 Audio Eng. Soc. Conv., London, May 2010.

    We have found evidence suggesting that for musically naïve participants, when selecting among similar renditions of the same musical fragment, psychoacoustic roughness is an influencing factor on preference. We designed an experiment to compare the acceptability of three different music fragments, rendered with three different intonations, and contrasted the results with those of isolated chords of the same fragment–intonation combinations.

  70. Julián Villegas and Michael Cohen. “Roughometer”: Realtime Roughness Calculation and Profiling. In Proc. 125 Audio Eng. Soc. Conv., San Francisco, October 2008.

    A software tool capable of determining auditory roughness in real-time is presented. This application, based on Pure-Data (Pd), calculates the roughness of audio streams using a spectral method originally proposed by Vassilakis. The processing speed is adequate for many realtime applications, and results indicate limited but significant agreement with an internet application of the chosen model. Finally, the usage of this tool is illustrated by the computation of a roughness profile of a musical composition that can be compared to its perceived patterns of ‘tension’ and ‘relaxation.’

  71. Mohammad Sabbir Alam, Michael Cohen, Ashir Ahmed, and Julián Villegas. Figurative Privacy Control of Sip-based Narrowcasting. In IEEE AINA: Int. Conf. on Advanced Information Networking and Applications, Gino wan, Japan, March 2008.

    In traditional conferencing systems, participants have little or no privacy, as their voices are by default shared with all others in a session. Such systems cannot offer participants the options of muting and deafening other members. The concept of narrowcasting can be applied to make these kinds of filters available in multimedia conferencing systems. Our system treats media sinks (in the simplest case, listeners) as full citizens, peers of the media sources (conversants’ voices), and we defined therefore duals of mute select: deafen attend, which respectively block a sink or focus on it to the exclusion of others. In this article, we describe our prototyped system, which uses existing standard Session Initiation Protocol (sip) methods to control fine-grained narrowcasting sessions. The design considers the policy configured by the participants and provides a policy evaluation algorithm for media mixing and delivery. We have integrated a ‘virtual reality’-style interface with this sip backend to display and control articulated narrowcasting with figurative avatars.

  72. Michael Cohen, Ishara Jaysingha, and Julián Villegas. Spin-Around: Phase-Locked Synchronized Rotation and Revolution in Multistandpoint Panoramic Browsers. In Proc. IEEE CIT 2007: 7th Int. Conf. on Computer and Information Technology, pages 511–516, Aizu-Wakamatsu, Japan, October 2007.

    Using multistandpoint panoramic browsers as displays, we have developed a control function that synchronizes revolution and rotation of a visual perspective around a designated point of regard in a virtual environment. The phase-locked orbit is uniquely determined by the focus and the start point, and the user can parameterize direction, step size, and cycle speed, and invoke an animated or single-stepped gesture. The images can be monoscopic or stereoscopic, and the rendering supports the usual scaling functions (zoom/unzoom). Additionally, via sibling clients that can directionalize realtime audio streams, spatialize hdd-resident audio files, or render rotation via a personal rotary motion platform, spatial sound and proprioceptive sensations can be synchronized with such gestures, providing complementary multimodal displays.

  73. Julián Villegas and Michael Cohen. Synœsthetic Music or the Ultimate Ocular Harpsichord. In Proc. IEEE CIT: 7th Int. Conf. on Computer and Information Technology, pages 523–527, Aizu-Wakamatsu, Japan, October 2007.

    We address the problem of visualizing microtuned scales and chords such that each representation is unique and therefore distinguishable. Using colors to represent the different pitches, we aim to capture aspects from the musical scale impossible to represent with numerical ratios. Inspired by the neurological phenomenon known as synæshesia, we built a system to reproduce microtuned midi sequences aurally and visually. This system can be related to Castel’s historic idea of the ‘Ocular Harpsichord.’

  74. Julián Villegas and Michael Cohen. Möbius Tones and Shepard Geometries: An Alternative Synæsthetic Analogy (poster). In Proc. SIGGRAPH NPAR: 5th Int. Symp. on Non-Photorealistic Animation and Rendering, San Diego, August 2007.

    “Why did the chicken cross the Möbius strip?” – “Because it wanted to get to the same side.” We created a 3d animation to illustrate a different visual analogy for Shepard tones based on the well-known “Möbius Strip II” by Escher. This animation presents a sphere that, like the ants in Escher’s woodcut, moves longitudinally over the surface. The path followed by the ball varies its transverse position randomly but smoothly. We sample this path to render a melody of Shepard tone dyads, each tone in the dyad having a frequency equivalent to the position of the ball relative to the edge of the surface. The idea of using a Möbius strip to create music is not new. Tremblay, for example, used it to show how to construct a music-box able to play sequences backward and forward. However, our work differs from other developments in the use of this non-orientable geometry to illustrate the mentioned aural paradox.

  75. Julián Villegas and Michael Cohen. Local Dissonance Minimization in Realtime. In Proc. ACM SIGMAP: Int. Conf. on Signal Processing and Multimedia Applications, Barcelona, July 2007.

    This article discusses the challenges of applying the tonotopic consonance theory to minimize the dissonance of concurrent sounds in real-time. It reviews previous solutions, proposes an alternative model, and presents a prototype programmed in Pd that aims to surmount the difficulties of prior solutions.

  76. Julián Villegas and Michael Cohen. Dsp-based Realtime Harmonic Stretching. In Proc. Hc-2006: Ninth Int. Conf. on Humans and Computers, pages 164–168, Aizu Wakamatsu, Japan, September 2006.

    This paper introduces a new way to express harmonic stretching in realtime based on Pd and overcoming the limitations imposed by the midi protocol in previous solutions. Applicable to non-percussive sounds.

  77. Julián Villegas and Michael Cohen. Melodic Stretching with the Helical Keyboard. In Proc. Enactive: 2nd Int. Conf. on Enactive Interfaces, Genoa, Italy, November 2005.

    A new technique to visually and aurally express melodic stretching in realtime based on the Helical Keyboard (a Java 3D application) and the MIDI protocol.

Patents

  1. Julián Villegas and Yan Pei. 関数特定方法、関数特定プログラム及び情報処理装置, Jul. 2026. Japanese patent application number 2026-178535 (Method for Identifying Functions, Program for Identifying Functions, and Information Processing Apparatus).
  2. Julián Villegas. 音声信号処理プログラム及び音声信号処理装置, Mar. 2023. Japanese patent number 2023-048058 (Speech signal processing program and speech signal processor).
  3. Julián Villegas and Anh T. Pham. 屋内位置特定システム、携帯端末及びコンピュータプログラム, Nov. 2021. Japanese patent number 6971467 (Indoor localization system using near-ultrasound signals).
  4. Julián Villegas. スピーカから再生される音の定位化方法、及びこれに用いる音像定位化装置, Sep. 2020. Japanese patent number 6770698 (Sound spatialization by equalizing filters and delay adjustments).

Books & chapters

  1. Yukki Baldoria, B. Paris Fleming, Mei Wan Rachel Liu, Julián Villegas, and Seunghun J Lee. ICU language database series 8: Comparative Zapotec. In ICU Working Papers in Linguistics (ICUWPL), number 22, pages 225–349. International Christian University, Tokyo, May 2022.
  2. Yuta Kariyado, Camilo Arévalo, and Julián Villegas. Auralization of three-dimensional cellular automata. In Romero et al., editor, Artificial Intelligence in Music, Sound, Art and Design, pages 161–170. Springer Int. Publishing, Apr. 2021. DOI: 10.1007/978-3-030-72914-1_11.

    An auralization tool for exploring three-dimensional cellular automata is presented. This proof-of-concept allows the creation of a sound field comprising individual sound events associated with each cell in a three-dimensional grid. Each sound-event is spatialized depending on the orientation of the listener relative to the three-dimensional model. Users can listen to all cells simultaneously or in sequential slices at will. Conceived to be used as an immersive Virtual Reality (VR) scene, this software application also works as a desktop application for environments where the VR infrastructure is missing. Subjective evaluations indicate that the proposed sonification increases the perceived quality and immersability of the system with respect to a visualization-only system. No subjective differences between the sequential or simultaneous presentations were found.

  3. Edward Ly and Julián Villegas. Genetic Reverb: Synthesizing Artificial Reverberant Fields Via Genetic Algorithms. In Romero et al., editor, Artificial Intelligence in Music, Sound, Art and Design, LNCS 12103, pages 90–103. Springer International Publishing, Switzerland, April 2020. DOI 10.1007/978-3-030-43859-3_7.

    We present Genetic Reverb, a user-friendly VST 2 audio effect plugin that performs convolution with Room Impulse Responses (RIRs) generated via a Genetic Algorithm (GA). The parameters of the plugin include some of the standard room acoustics parameters mapped to perceptual correlates (decay time, intimacy, clarity, warmth, among others). These parameters provide the user with some control over the resulting RIRs as they determine the fitness values of potential RIRs. In the GA, these RIRs are initially generated via a Gaussian noise method, and then evolved via truncation selection, multi-point crossover, zero-value mutation, and Gaussian mutation. These operations repeat until a certain number of generations has passed or the fitness value reaches a threshold. Either way, the best-fit RIR is returned. The user can also generate two different RIRs simultaneously, and assign each of them to the left and right stereo channels for a binaural reverberation effect. With Genetic Reverb, the user can generate and store new RIRs that represent virtual rooms, some of which may even be impossible to replicate in the physical world. An original musical composition using the Genetic Reverb plugin is presented to demonstrate its applications. (The source code and link to the demo track is available at https://github.com/edward-ly/GeneticReverb).

  4. Michael Cohen and Julián Villegas. Applications of audio augmented reality. Wearware, everyware, anyware, and awareware. In Fundamentals of Wearable Computers and Augmented Reality, chapter 13, pages 309–329. CRC Press, 2nd edition, July 2015. DOI 10.1201/b18703-17.
  5. Michael Cohen, Julián Villegas, and Woodrow Barfield. Special issue on spatial sound in virtual, augmented, and mixed-reality environments, volume 19. Springer-Verlag, 2015. DOI: 10.1007/s10055-015-0279-z.
  6. Julián Villegas. Psychoacoustic Roughness Applications in Music: On Automatic Retuning and Binaural Perception. LAP LAMBERT Academic Publishing, 2012.
  7. Julián Villegas and Michael Cohen. Mapping Musical Scales Onto Virtual 3D Spaces. In Yôiti Suzuki, Douglas Brungart, Hiroaki Kato, Kazuhiro Iida, Densil Cabrera, and Yukio Iwaya, editors, Principles and Applications of Spatial Hearing. World Scientific, 2011. DOI 10.1142/9789814299312_0036.

    We introduce an enhancement in the Helical Keyboard, an interactive installation displaying three-dimensional musical scales aurally and visually. This improvement in the audio display is intended to facilitate didactic purposes by enhancing users’ immersion in a virtual environment. The new system allows spatialization of audio sources with elevation angles between −40° and +90° and azimuth angles between 0° and 355°. In this fashion, we could overcome previous limitations on the audio display of the Helical Keyboard, for which we heretofore usually displayed only azimuth.

  8. Sabbir Alam, Michael Cohen, Julián Villegas, and Ashir Ahmed. Narrowcasting in SIP: Articulated Privacy Control. In Syed Ahson and Mohammad Ilyas, editors, SIP Handbook: Services, Technologies, and Security of Session Initiation Protocol, chapter 14, pages 323–345. CRC Press, 2009.
  9. Felipe Millán Constaín, Juan Camilo Paz, Alfredo Roa, Julián Villegas, Nicolás Carranza, Diego Briceño, and Alex Mera. Medición de la Productividad del Valor Agregado (Added Value Productivity Measurement). Servicio Nacional de Aprendizaje (Sena), 2nd edition, 2003. in Spanish. [Online; accessed 31-Aug-2008]: http://cnp.org.co.

Invited articles

  1. Julián Villegas. 音程感覚の習得とより良い歌唱体験の補助 (improving singing experience for people with tuning difficulties). Sound, 32:10–14, Jan 2017. In Japanese.

    Based on the association of psychoacoustic roughness and musical pitch, and inspired by the common tuning technique of eliminating aural beats between the strings of an instrument, we hypothesize that users of our system could adjust their intonation in order to minimize the interference between their current and desired pitch (a modulated version of his or her current voice). It is our hope that this process could lead to long-term singing improvements, as well. This work-in-progress report discusses implementation issues, expression possibilities, and future evaluations of this tool as an alternative for improving singing skills in self-refrained singers.

  2. Michael Cohen, Julián Villegas, and Woodrow Barfield. Special issue on spatial sound in virtual, augmented, and mixed-reality environments, volume 19. Springer-Verlag, 2015. DOI: 10.1007/s10055-015-0279-z.
  3. Julián Villegas and Michael Cohen. 音楽を探る オペレーションズ・リサーチ的手法を使って(exploring tonal music through operational research methodology). Communications of the Operations Research Society of Japan, 54(9):554–562, October 2009. In Japanese.

    Two operational research applications in music are presented. Initially, the mapping of musical scales into multi-dimensional topologies is discussed, and the advantages of projecting these structures into simple spaces explained. We also present the Helical Keyboard, an interactive installation displaying three-dimensional musical scales aurally and visually. Subsequently, the problem of minimizing musical dissonance between audio streams in realtime is discussed, and a solution based on local minima search described.

  4. Julián Villegas, Yuuta Kawano, and Michael Cohen. Harmonic Stretching with the Helical Keyboard. 3D Forum: J. of Three-Dimensional Images, 20(1):29–34, 2006.

    An extended version of the paper published in Proc. HC-2005: Eighth International Conference on Humans and Computers, introducing other possibilities to achieve harmonic stretching using only the midi protocol.

Invited talks

  1. Julián Villegas. Analysis of fundamental frequency trajectories using generalized additive models. In 6th Int. Symp. on Applied Phonetics (ISAPh2026), University of Hiroshima, Sep. 2026. https://julovi.codeberg.page/gam_tutorial.
  2. Julián Villegas. Compression and personalization of head-related transfer functions using machine learning. Joint International Summer School Sofia University – University of Aizu, Aug. 2025.
  3. Julián Villegas and Seunghun J. Lee. Phonetics of the laryngeal contrast: an overview of data collection methods. 1st meeting of the ILCAA Joint Research Project “Phonetic typology from cross-linguistic perspectives”, Jul. 2021.

    This talk presents an overview of data collection methods in studies that investigate the laryngeal contrast in various languages in the past years. The majority of studies analyze acoustic data with correlates such as VOT, closure duration, phonation. Analyses based on articulatory data such as Electroglottography (EGG) and Ultrasound are also employed to report the larynx movement concerning the voicing contrast. We also share analytical tools such as EGGNOG that can be used to process and report articulatory data.

  4. Julián Villegas. Current trends of spatial sound in virtual reality. 2nd meeting of the ASJ Tohoku region, Nov 2019.
  5. Seunghun J. Lee and Julián Villegas. Correlation analysis of microvariation data by R (programming language). First meeting of ILCAA Joint Research Project “Typological Study of Microvariation in Bantu (2)”, October 2019.
  6. Julián Villegas. Introduction to statistical analysis of phonetic data in R. In 2nd Int. Symp. on Applied Phonetics (ISAPh2018), University of Aizu, Sep 2018. DOI 10.21437/ISAPh.2018-5.

    This manuscript summarizes the main points discussed during a workshop with the same title presented during the 2nd International Symposium on Applied Phonetics—ISAPh2018. Because of the time constraints of the workshop, this is a very limited introduction to R created with the intention of showing how to use R in the analysis of phonetic data and encourage attendees to use it in their research. Dataset for the examples, for a do-it-yourself part, and a tutorial are freely available from http://onkyo.u-aizu.ac.jp/classes/ez.

  7. Julián Villegas. Data extraction of audio and EGG recordings. In Proc. 3.5 Joint Workshop of the Phonetics Phonology and New Orthographies (PhoPhoNo), Universität Bern (Switzerland), March 2018.
  8. Julián Villegas. Automatic prediction of creaky voice with psychoacoustic roughness. In N/A, Centre National de la Recherche Scientifique, Laboratoire Psychologie de la Perception, (Paris), Sep 2017.
  9. Julián Villegas. Measuring acoustic features from audio and EGG recordings. In Proc. Int. Electroglottography Workshop, International Christian University (Tokyo), Oct 2016.
  10. Julián Villegas. Parallels between musical consonance and speech intelligibility. In Summer school of linguistics, Dac̆ice, Czech Republic, Aug. 2012.

    In this talk, the origins of speech intelligibility and musical consonance are discussed. Physical, perceptual, and cognitive causes of these complex phenomena have been identified and the understanding of their interactions is still a very active topic of research. We will focus on acoustic and psycho-acoustic features of speech and music and present the effect of them on intelligibility and consonance. Particularly, the role of fundamental frequency in both phenomena will be discussed: we will show the effect of scaling the fundamental frequency of voiced regions on objective speech intelligibility. The frequency scaling is driven by maximizing the glimpse proportion, inspired by musical consonance maximization techniques.

  11. Julián Villegas. The role of speech rate of speech intelligibility. In Summer school on linguistics, Dac̆ice, Czech Republic, Aug. 2012.

    Speech rate has been identified as one of the main differences between speaking styles. Clear speech (or the speech produced by someone who has been asked to speak clearly) and Lombard speech (or speech produced in noisy environments) are more intelligible than other styles (like casual and ”normal” speech), and also have a slower speech rate. In this talk, we will discuss the role of speech rate on intelligibility, show how to artificially modify speech rate either by stretching or compressing an utterance or by time-aligning one to another. Rather than discussing the inner mechanisms of the signal processing, we will focus on understanding the differences between the two modalities and how to use existing software applications to modify duration.

  12. Julián Villegas. Acoustic modifications of speech and intelligibility in noise. In Proc. Int. Symp. on Spatial Media (ISSM), Aizu-Wakamatsu, Japan, Mar. 2012.
  13. Julián Villegas. Listen to what i say: environment-aware speech production. In Japan-Woche der FH Düsseldorf Wkshp. on Mixed Reality and Virtual Environments, Düsseldorf, May 2011.

    Major challenges to adapt all forms of speech output to a given auditory context (e.g., noisy or highly reverberant environments, second language or hearing-impaired listeners, etc.) based on human speaker strategies are discussed. Ongoing research aimed at increasing speech intelligibility in real-time without compromising speech quality (or fatiguing the listener) is described, and software applications used in this research are presented. This talk will also present auditory demonstrations of natural and artificial speech modifications.

  14. Julián Villegas. Unconventional 3D-sound Controllers. In Michael Cohen, editor, Proc. Int. Symp. on Spatial Media (ISSM), Aizu Wakamatsu, Japan, March 2011.

    Wireless technologies allow the introduction of sensors in otherwise unanticipated devices, increasing the opportunities of interaction with real-world, daily-life things. In this paper the idea of controlling three-dimensional audio by means of flying discs is presented, the prototype implementation (based on gyroscopes, Xbee radios, Arduino micro-controllers, Pure-Data and Quartz Composer patches) helps to understand the challenges, capabilities and limitations of the underlying technologies that make such interactions possible.

  15. Michael Cohen and Julián Villegas. Spatial Sound and Entertainment Computing. In Proc. Int. Conf. on Entertainment Computing (ICEC), Seoul, September 2010.

    This tutorial introduces the theory and practice of spatial sound for entertainment computing, including psychophysical (psychoacoustic) basis of spatial hearing, outlines the mechanism for creating and displaying spatial sound the hardware and software used to realize such systems, display configurations, and reviews some applications of spatial sound to entertainment computing, especially multimodal interfaces, featuring spatial sound. Many case studies reify the explanations; animations, videos, and live demonstrations are featured

Software

  1. Camilo Arevalo, Naoki Fukasawa, and Julián Villegas. Non-personalized HRIR databases with and without floor reflections, June 2021. DOI: 10.5281/zenodo.4954711.

    Non-personalized HRIR databases in SOFA format with and without floor reflections. Floor reflections were simulated with a plywood board between a Head-And-Torso Simulator (HATS) and a dodecahedral loudspeaker. These recordings were captured at the anechoic chamber of the University of Aizu.

  2. Julián Villegas. Eggnog: Electroglottography not obligatory guidance, 2020. Available [September 29, 2026] from https://bitbucket.org/julovi/eggnog/.

    Eggnog: ElectroGlottoGraphy Needs Optional Guidance It requires to have the Signal Processing Toolbox installed in Matlab.

  3. Julián Villegas. Roughometer, 2020. Available [September 29, 2026] from https://bitbucket.org/julovi/roughometer.

    Real-time Roughness Calculation and Profiling. A Pure-data tool capable of determining auditory roughness in real-time. This application calculates the roughness of audio streams using a spectral method originally proposed by Vassilakis. Please cite: J. Villegas and M. Cohen. “Roughometer”: Realtime Roughness Calculation and Profiling. In Proc. 125 Audio Eng. Soc. Conv., San Francisco, October 2008.

  4. Julián Villegas. Shifter, 2020. Available [September 29, 2026] from https://bitbucket.org/julovi/shifter.

    A PSOLA pitch shifter in Pure-data

  5. Julián Villegas. Self-duet, 2018. Available [September 29, 2026] from https://bitbucket.org/julovi/self-duet.

    Self-duet is a proof-of-concept of Computer-assisted singing experience created by Julián Villegas in collaboration with Mutsuko Ishihara. It comprises a DSP program to be run in a server, and a GUI to be run in a tablet or smartphone running TouchOSC (https://hexler.net/software/touchosc). Recordings of the results, and more information are available at http://onkyo.u-aizu.ac.jp/software/self-duet/ Cite as: Julián Villegas. 音程感覚の習得とより良い歌唱体験の補助 (improving singing experience for people with tuning difficulties). Sound, 32:10–14, Jan 2017. In Japanese.

  6. Julián Villegas. Creakiness by roughness, 2017. Retrieved September 29, 2026. Available from https://bitbucket.org/julovi/creakbyroughness.

    Matlab routine to automatically predict creakiness using a model of psychoacoustic roughness

  7. Julián Villegas. Beating and Roughness. The Wolfram Demonstrations Project, September 2010. [Online; accessed 23-Sep-2010]: http://demonstrations.wolfram.com/BeatingAndRoughness.

    A demonstration of beating sinusoids showing fluctuation strength, roughness, and tone separation.

Others

  1. Julián Villegas. Pliegos de Cordel. Lulu.com, 2024. ISBN: 978-1-304-18514-3.
  2. Julián Villegas. The 60th anniversary of Tohoku chapter of the Acoustical Society of Japan. chapter Computer Arts Lab at the University of Aizu. Acoustical Society of Japan, 2016.
  3. Julián Villegas. Psychoacoustic Roughness Applications in Music: On Automatic Retuning and Binaural Perception. PhD thesis, University of Aizu, Aizu-Wakamatsu, Japan, March 2010.

    The goal of this study is to help to understand the influence of psychoacoustic roughness in music. Roughness is an auditory attribute produced by rapid temporal envelope fluctuations (normally resulting from wave interferences), and it has been related to musical dissonance. After reviewing the main theories that explain the origin of roughness, a software program created for the purpose of this research is presented. This software application, based on a spectral model to predict roughness (a physical predictor of the auditory attribute), is able to control (usually, to reduce) the predicted roughness of a sound ensemble in realtime. Experimental results were analyzed with a standard measurement software tool to corroborate the predicted roughness reduction. The audio output of this software application was compared by human subjects with renditions of the same musical content using some well known tuning systems (twelve tones equal tempered and just tuning). The results of these subjective experiments are presented and analyzed. Preliminary results on binaural roughness perception are presented at the end of the dissertation as a new direction of research. Contributions of the present work include the creation of an adaptive tuning program capable of retuning audio streams in realtime to minimize the measured roughness due to the interaction between sounds (extrinsic roughness). To the best of our knowledge, this procedure had been applied only to midi sequences for which realtime constraints implied oversimplifications that are not assumed in our program. We were able to determine that roughness by itself can explain musical preference among musically naïve participants. In the analysis of that experiment, we found that, contrary to popular belief, predicted roughness of 12-TET intervals is not always greater than pure intervals. This discovery correlates with preference choices reported by participants. We also show that current roughness models need to be revised to include the effect of binaural cues. Other minor contributions include several entries in Wikipedia (e.g., Vicentino’s keyboard layout) and to Mutopia (e.g., Bach choral BWV 264).

  4. Julián Villegas. Encuentros entre Colombia y Japón: homenaje a 100 años de amistad. chapter De como el mundo es un pañuelo y de las misteriosas maneras (Of how the world is a handkerchief and the mysterious ways). Colombian Ministry of Foreign Affairs, Bogotá D.C., Colombia, 2010. (Fiction, in Spanish).
  5. Julián Villegas. Local Consonance Maximization in Realtime. Master’s thesis, University of Aizu, Aizu-Wakamatsu, Japan, September 2006.

    Although the problem of maximizing consonance in tonal music has been addressed before, every solution reflecting the technological advances of its epoch, and considering that current theories to explain this psychoacoustical phenomenon are generally satisfactory, there are still vast unexplored aspects of this area, since even most recent solutions lack adequate mechanisms to apply such techniques in realtime scenarios. In general, the most advanced achievements in this field are based on the midi protocol for controlling the pitch of simultaneous notes, inheriting the protocol limitations in terms of dependency on the quality of the synthesizer for satisfactory results, scalability, accuracy, veracity, etc. Besides that, timbres are generally known a priori for these techniques, so their application to unknown timbres requires digitization and analysis of sound samples, making such techniques unsuitable for realtime situations. This thesis summarizes the main theories about consonance and its relation to musical scales, reviews several previous solutions as well as the state of the art, proposes an alternative model to adaptively adjust consonance in a polyphonic scenario based on the tonotopic dissonance paradigm (presented by Plomp and Levelt, having been previously developed by Sethares), and presents a prototype of this model that aims to surmount the difficulties of prior solutions by performing realtime analysis and pitch adjustment programmed in Pure-data (Pd), a data flow DSP environment for realtime audio applications. The results are analyzed to determine the efficacy and efficiency of the proposed solution.

  6. Julián Villegas. Diseño e Implementación de un Algoritmo Genético para la Asignación de Aulas en la Universidad del Valle (Design and Implementation of a Genetic Algorithm for Timetabling at the University of Valle). Undergraduate honors thesis, University of Valle, Cali, 2001. In Spanish.

    Cómo aplicar las ventajas de las técnicas de procesamiento paralelo y de inteligencia artificial, específicamente de Algoritmos Genéticos, en la solución del problema de asignación de aulas, inicialmente en la Universidad del Valle, además de comparar la eficiencia del sistema actual con el nuevo esquema propuesto, es el objeto de esta tesis. En el primer capítulo se presenta una Introducción a la Teoría de la Computación en Paralelo: se muestran las diferentes fuentes de paralelismo en un programa y las diferentes arquitecturas de software y hardware disponibles para lograr paralelismo. Se comparan las diferentes opciones con las necesidades de la presente tesis. En el segundo capítulo se presenta una Introducción General a los Algoritmos Genéticos: se explica su funcionamiento, se presentan los operadores genéticos más comunes, las estrategias de selección de más aceptación; también, se muestra la relación que hay entre la computación en paralelo y los algoritmos genéticos. Finalmente, se hace una presentación de una implementación popular de un algoritmo genético paralelo. En el tercer capítulo, se presenta el problema general de asignación de aulas, de horarios, y sus principales características. En el cuarto capítulo se presenta la evolución del problema de la asignación de aulas y horarios y el estado del arte. En el quinto capítulo se analiza el proceso actual (enero – mayo 2000) de asignación de aulas en la Universidad del Valle; se determinan las entradas y salidas del proceso, se calcula la dimensión del problema, se recopilan los requerimientos de los entes involucrados, y se identifican las restricciones y prioridades tenidas en cuenta en el proceso de asignación de aulas. En el sexto capítulo se propone una solución al problema de asignación de aulas basada en algoritmos genéticos paralelos: resume el proceso de instalación de 6 PVM, la configuración empleada y algunos detalles que no son tan claros en la documentación que viene con las fuentes de PVM; además, se muestra la instalación de SSH y la manera de emplearlo conjuntamente con PVM en una red donde la seguridad es importante. En el séptimo capítulo, se hace un análisis de los resultados obtenidos y se discuten posibles desarrollos posteriores, mejoramientos y refinamientos del algoritmo genético propuesto. En el octavo se anexan el código fuente en C de los programas que hacen parte de la solución (el código del algoritmo genético maestro y el esclavo), como también los códigos SQL de las consultas realizadas a la base de datos para extraer la información necesaria para la asignación. Además, un papel presentado en GECCO – 2000 (Genetic and Evolutive Computation Conference - 2000). En el noveno capítulo se presentan las conclusiones generales del presente proyecto de grado. El décimo capitulo presenta las referencias bibliográficas empleadas en el desarrollo del presente proyecto.

Original music

  1. Julián Villegas. GoldenM. 360 degrees of 60×60 event at icmc 2010: The Int. Computer Music Conf., June 2010.

    GoldenM is a computer-generated composition created in Pd. It is an arrangement for three voices and filtered white noise. In GoldenM, each voice has a spectrum based on the golden ratio (about 1.61803), and the pitch set was selected using the minima of the dissonance function as proposed by Vassilakis. Rhythm and the spatialization are generated using Markov chains. The purpose of the composition is to use the golden ratio, in unnatural ways preserving, in some extent, its esthetic nuance. The result, at moderate volume, resembles (at least to the author) the sound of chimes and bells used in Asian musical traditions.

  2. Julián Villegas. Original Music for “El Proyecto del Diablo”. Tv Broadcasted by Rostros y Rastros (Uvtv), 1999. Documentary directed by Óscar Campo (25 Min.). www.imdb.com/title/tt0483127.

    “Todos los caminos me llevan al infierno… Pero si el infierno soy yo!”, es la frase con la que inicia el monólogo que narrará a través de todo el documental La Larva, quien no sabemos si está vivo o muerto. En El Proyecto del Diablo conocemos la historia y los pensamientos más reveladores de este hombre: Desde que probó la marihuana en el colegio, pasando por sus experiencias tropeleras en la universidad, hasta los momentos de mayor éxtasis con las drogas y su estadía en la cárcel por doce años. Descubrimos así la mentalidad de quien se dejó atrapar y llevar por la maldad.

  3. Julián Villegas. Original Music for “Aquí No Canta Nadie”. Tv Broadcasted by Rostros y Rastros (Uvtv), 1998. Short Film directed by Pilar Chávez (15 min.).

    Esta es una de las tantas historias que puede acontecer el campo Colombiano. Una noche, en una finca, tres niños escuchan los relatos de terror narradas por una mujer, quien les advierte que no deben salir de casa, pues el diablo puede llevarlos a una muerte horrible. Al amanecer uno de los niños ha desaparecido; sus dos hermanos lo buscan por toda la casa sin encontrarlo por ninguna parte. El hermano mayor se atreve a salir en su búsqueda y logra regresar con su hermano a casa, sin embargo, en esta travesía se confrontaron sus recuerdos con una cruda verdad.

  4. Julián Villegas. Original Music for “El Terminal”. Tv Broadcasted by Rostros y Rastros (Uvtv), 1998. Documentary directed by Margarita Arbeláez, Luz Elena Luna, Ximena Bedoya, Andrea Rosales, Juan Camilo Duque, Claudia Villegas, and María Fernanda Gutiérrez (26 min.).

    Este documental retrata la rutinas, los personajes y sucesos que acontecen en El Terminal de Transportes de Cali. El paso del tiempo en el documental transcurre al ritmo de la espera y la ansiedad de los viajeros: Gente que va y viene; otros se despiden y se alejan; otros llegan sin conocer. El día pasa y llega la noche, lo que marca diferentes ritmos de la vida en El Terminal.

  5. Julián Villegas. Original Music for “El Ojo de Buziraco”. Tv Broadcasted by Rostros y Rastros (Uvtv), 1997. Four Short Films directed by David Bohórquez (26,25,26,24 min.).

    A four part series, granted by the Ministry of Culture of Colombia, about the myths and legends of the Colombian Pacific coast which draws a thin line between fact and fiction. El ojo de Buziraco I: Nadie vio nada El Ojo de Buziraco es una serie compuesta por cuatro capítulos que reactualiza algunos mitos colombianos (mantenidos a través de la tradición oral) en ámbitos urbanos. En el primer capitulo titulado Nadie vio nada, dos hombres de edad madura reconstruyen, a través de su testimonio, la historia que oyeron de sus antecesores a cerca del mito del Mohan, mounstro de los ríos; y paralelamente, se desarrolla una historia de ficción en la que una banda criminal de la ciudad se verá acechada y atacada por este personaje. El ojo de Buziraco II: Tente en el aire En el segundo capítulo de la serie El Ojo de Buziraco, la historia gira en torno a un grupo de universitarios, estudiantes de audiovisuales, que indagan acerca de la función de los “cuentos de miedo” en la educación de tres adultos mayores, a quienes dichas historias fueron contadas por sus padres y abuelos. La novia de uno de los jóvenes ha fallecido recientemente y a medida que los testimonios de las entrevistas describen el mito de La Tunda, él comienza a experimentar acerca- mientos con la difunta: la ve en pesadillas recurrentes, en la pantalla de los monitores de edición y a través de la cámara; al final, la presencia fantasmal aparece y lo convierte en una víctima más de La Tunda. El ojo de Buziraco III: La guerra de Mandrágora La guerra de Mandrágora es el tercer capítulo de la serie El Ojo de Buziraco donde los entrevistados sostienen que las brujas sí existen. En la puesta en escena, una mujer acude a una bruja para buscar solución a las pesadillas que atormentan a su marido; luego, la mujer en su soledad y frente a la sospecha de ser engañada por su esposo, vuelve donde la bruja quien ha decidido que ella debe ser su sucesora. Una vez terminados los ritos de iniciación, aparece el diablo castigando a la bruja y la historia tiene un final inesperado. El ojo de Buziraco IV: El Vampiríparo El último capítulo de la serie El Ojo de Buziraco, acontece en un salón de clases donde el estudiante más «atontado», Medardo, comienza a tener una serie de alucinaciones. En los momentos de mayor presión y frustración, se ve a si mismo como un nosferatu, siempre en la búsqueda de un mentor: un vampiro «real». Dicha obsesión, llevará a Medardo a vivir múltiples situaciones en las que se verá comprometida su integridad. Entre tanto, los entrevistados, que son la cuota documental del audiovisual, sostendrán en sus testimonios la no-existencia de los vampiros y recrearán historias que se convirtieron en mitos urbanos y que posiblemente inspiraron la leyenda acerca de la existencia del vampirismo en la ciudad de Cali.

  6. Julián Villegas. Original Music for “Mario y las Voces”. Tv Broadcasted by Rostros y Rastros (Uvtv), 1997. Short Film directed by Carlos Espinosa (15 min.).

    Relato futurista en el que las máquinas ejercen pleno control sobre los hombres. Esta tecno-sociedad de la represión, sin embargo, oculta una verdad que le es revelada a Mario desde una suerte de submundo habitado por quienes han decidido resistir: tras el poder de las máquinas se esconde el poder de algunos pocos hombres.

  7. Julián Villegas. Original Music for “No, no…baby”. Tv Broadcasted by Rostros y Rastros (Uvtv), 1997. Short Film directed by Diego Pérez (13 min.).

    No, no…baby es una historia que combina elementos de drama y suspenso alrededor de las extrañas, violentas y perversas relaciones que se tejen en un grupo de amigos, entregados con pasión al consumo de la carne. Los excesos desatan sus instintos caníbales al momento que aparecen los conflictos sentimentales entre ellos. El desenlace de esta historia es llevado al extremo cuando vemos que el grupo termina devorándose mutuamente mientras observan películas de Stanley Kubrick.

Last updated: September 29, 2026