Open Reader

While my guitar gently speaks — Todd Fisher, Philo Ventures

completed 18:35 Aug 18, 2026 Watch on YouTube

Current Status

completed

Video ID

E_Txocq-Lrw

RAG / Chat

Enabled
While my guitar gently speaks — Todd Fisher, Philo Ventures
Description

Someone in the audience asked the guitar what reality is, and the guitar answered. Todd Fisher's build routes a microphone through speech recognition into a local model and pushes the reply back out through the strings, which is the most recent step in a project that began with a much simpler question: how hard could it be to make a guitar speak? The lineage he draws runs from a pickup and an amplifier, through stomp boxes, to Peter Frampton sending guitar sound down a physical hose into his mouth. His own version is a plugin built with JUCE that drops into a DAW like any other effect. Getting it to say one word was straightforward. Getting it to say several meant slicing synthesized speech into words automatically, and that turned out to be the hard part. Energy gap segmentation cuts wherever the signal falls toward silence, which fails because running speech often has no silence between words at all. A sonority peak syllabifier looks for vowels instead. Combining the two got close enough that he finished by dragging segment boundaries by hand. Singing needed a different stack again: the YIN algorithm to pull a fundamental frequency off each fretted note, a synthesized tone shaped by an envelope, then a vocoder, with pitch shifted samples from an open singing dataset baked ahead of time because the processing is far too heavy to run live. He also declines to play the song his title alludes to, on the grounds that this recording was going online. Speaker info: - https://www.linkedin.com/in/todd-b-fisher Timestamps: 0:00 - Live performances that stayed with him 2:44 - The guitar's evolution, up to the talk box 4:24 - A Halloween project on a garage door 6:06 - Building it as a JUCE plugin, and saying one word 7:50 - Slicing speech into words, and why that is hard 10:23 - Pitch detection with the YIN algorithm 11:15 - Synthesis, vocoder, and jamming with it 13:28 - A guitar that answers questions from the room 16:01 - Pitch shifted samples, and getting closer to si

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: Todd Fisher demonstrates a guitar-to-voice/singing prototype and uses it to argue that AI has made ambitious personal hardware-software projects materially easier to explore.
  • Why it matters: The talk offers a concrete example of assembling local AI, audio DSP, and a conventional plugin workflow into an interactive creative interface, while exposing the practical limits that still prevent a polished real-time product.
  • Best use: Use it as lightweight inspiration and a component map for voice-driven musical interfaces, not as a production-ready technical tutorial.

Executive Summary

Fisher frames the project around memorable live-performance effects, from Slipknot's rotating drum set to the stagecraft in Stranger Things: The First Shadow. His premise is that technology can create similarly surprising artistic experiences, and that current AI tools lower the barrier for engineers to pursue the side projects that have long sat on their idea lists.

His project began with a Stranger Things-themed Halloween installation: guitar notes controlled alphabet lights projected onto his garage door. He then pursued a more ambitious extension—making a guitar produce spoken and eventually sung language through a custom Logic Pro plugin.

The spoken system converts text to speech, segments the output into playable units, and triggers those units from guitar input. Automatic word segmentation proved unreliable because natural speech does not always contain silent gaps, so Fisher combined energy-gap detection with syllable-oriented analysis but ultimately retained manual segment editing.

For conversation, the prototype chains microphone speech-to-text through Whisper, a local LLM, and text-to-speech output played by the guitar. For singing, it detects guitar pitch with YIN and uses preprocessed vocal samples pitch-shifted with WORLD; this works only through offline pre-baking and is still limited to simple vowel sounds. The project is a credible proof of concept, but the live demo's glitches and manual/offline dependencies show that it remains exploratory.

Key Takeaways

  • Claim: AI can make personally meaningful experimental builds more accessible, but it does not eliminate engineering caveats or iteration. | Evidence: Fisher says the previous six months of AI progress helped push dormant personal projects forward, while explicitly qualifying that AI makes projects easier rather than frictionless. | Implication: Ken should treat AI as an accelerator for fast prototype assembly and exploration, not as a substitute for designing the difficult real-time, reliability, and UX layers. | Caveat: The remainder of the talk demonstrates substantial unresolved work: live-demo failures, manual speech segmentation, and offline vocal processing.
  • Claim: The most workable initial design is a DAW-native plugin that treats text-to-speech clips as musical assets triggered by guitar events. | Evidence: Fisher built a plugin for Logic Pro using audio tooling including JUCE, generated speech with Piper and Apple text-to-speech options, and first demonstrated a single spoken clip triggered from the guitar. | Implication: For creative-AI interfaces, embedding inside an established host environment can avoid building audio routing, effects chains, and performance controls from scratch.
  • Claim: Naively segmenting synthesized speech by silence is inadequate for turning continuous language into individually playable words. | Evidence: His energy-gap segmentation approach cuts waveform regions when decibel levels approach zero, but he notes that adjacent spoken words often have no usable silent interval. He added a Sonority Peak Syllabifier to identify vowel-led syllabic structure, then still exposed draggable manual segment boundaries. | Implication: Any agent or interface that must convert generated speech into controllable musical or interaction units needs human-editable alignment or a stronger forced-alignment pipeline rather than assuming waveform thresholding is reliable. | Caveat: The talk does not establish that the combined segmentation method is robust across speakers, languages, speaking styles, or arbitrary text.
  • Claim: A guitar can function as a conversational output device by chaining local speech recognition, local language-model inference, and voice playback. | Evidence: Fisher describes speaking into a microphone, transcribing with Whisper, sending the text to a local model, and placing the model's response onto the guitar; in the demo, the audience asks, "What is reality?" and the system replies with a guitar-voiced question. | Implication: The architecture is feasible as a local, privacy-preserving interactive loop, but it should be evaluated as a latency-sensitive orchestration system before being used in a live experience. | Caveat: The response is audibly choppy, the system has live-demo glitches, and Fisher provides no latency, model, hardware, or reliability measurements.
  • Claim: Making the guitar sound like it sings requires pitch-aware vocal synthesis rather than simply replaying spoken audio. | Evidence: Fisher uses the YIN pitch-detection algorithm to derive a guitar note, creates a synthesized carrier signal, and routes it through voice material in a talk-box/vocoder-inspired approach. He also describes balancing synthetic tone against the source voice with a clarity control. | Implication: Pitch extraction and carrier/source blending are the key controllable parameters for an expressive instrument-voice interface; natural vocal realism is a separate, harder synthesis problem. | Caveat: The showcased result remains closer to stylized vocoding/speech than natural singing.
  • Claim: The current singing implementation trades real-time responsiveness for feasible computation by precomputing pitch-shifted vocal samples. | Evidence: Fisher draws from VocalSet recordings and uses WORLD for pitch shifting, then maps pre-baked samples to guitar frets/notes. He says the process is too computationally heavy to perform live and therefore begins only with samples of the five vowel sounds. | Implication: For a live system, Ken should distinguish between pre-rendered asset playback and true generative audio synthesis; the former is dependable for demos, while the latter requires a dedicated real-time performance budget. | Caveat: Vowel-only samples cannot express intelligible lyrics or full vocal articulation, and pre-baking constrains improvisational content.

Detailed Brief

Creative lineage and interaction design

  • Claims: Fisher positions the project as part of the guitar's historical expansion from acoustic instrument to amplified rock instrument, pedal-driven effects platform, talk box, and software-defined production tool.; The intended value is experiential rather than utilitarian: a guitar that speaks or responds should create the same kind of surprise associated with memorable live-stage effects.
  • Evidence: He references the talk box associated with Peter Frampton, where guitar sound is physically shaped by the player's mouth.; His precursor Halloween app allowed typed custom messages, including "Happy Halloween," to be displayed through a Stranger Things-style alphabet light system as he played.; He compares a DAW to an IDE for musicians and producers, emphasizing the plugin as a familiar extension point rather than a standalone instrument platform.
  • Caveats: The presentation is primarily a personal project narrative and does not provide comparative evaluation against commercial vocal instruments, MIDI-guitar systems, or modern singing-synthesis products.
  • Implications: The strongest product insight is not merely "AI guitar" but designing a legible performance interaction: audiences should understand that an instrumental gesture is causing an unexpected linguistic or visual response.; Existing creative hosts and cultural references can make experimental interfaces immediately understandable to users and audiences.

Notable Concepts & Terms

  • JUCE: A framework Fisher recommends for building audio software; it serves as the likely application/plugin foundation for integrating DSP and DAW workflows.
  • DAW: Digital Audio Workstation, such as Logic Pro; Fisher characterizes it as an IDE-like environment for musicians, making it the host for his custom plugin.
  • Piper: A text-to-speech option Fisher explored for converting typed or generated text into audio clips that the guitar can trigger.
  • Energy gap segmentation: A waveform-based method that identifies likely word boundaries from low-volume gaps; useful as a first pass but unreliable for continuous speech.
  • Sonority Peak Syllabifier: A syllable-identification approach based on speech sonority and vowel structure, used to supplement silence-based segmentation.
  • YIN pitch algorithm: The pitch-detection method used to derive the guitar's fundamental note so vocal material can be mapped or synthesized at the played pitch.
  • Whisper: The speech-to-text component in the microphone-to-local-LLM-to-guitar conversation pipeline.
  • VocalSet and WORLD: VocalSet supplies recorded singer material; WORLD is used to pitch-shift and prepare it for note-mapped guitar playback, currently offline rather than live.

Operator Notes / Why Ken Should Care

  • If exploring voice-enabled instruments or physical AI interfaces, prototype first as a plugin within an existing host such as a DAW rather than building standalone audio infrastructure.
  • Require end-to-end latency, CPU, failure-rate, and recovery measurements before treating a local Whisper-to-LLM-to-audio loop as viable for live interaction.
  • Keep manual alignment/editing controls in any speech-to-performance workflow until automated segmentation is validated on the actual target voices and language set.
  • Decide early whether the product requirement is expressive prerecorded vocal playback or real-time generative singing; they impose fundamentally different compute and content pipelines.
  • Avoid using this talk as technical due diligence for audio synthesis quality: it offers architecture clues and failure modes, but no benchmarked implementation details.

Source/Metadata

  • Title: While my guitar gently speaks — Todd Fisher, Philo Ventures
  • Transcript words: 2942
  • Duration seconds: 1115
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript.

Transcript

2918 words en Processed in 138.4s

So I am Todd Fisher. I love the guitar. That's one of my passions in life. Today I want to talk about this project I've been working on for a while here, effectively making my guitar speak. But of course I want to start out framing it under this awesome premise. We've all been to live performances where your mind was just blown. It was awesome. The first one I remember, way back, I was in high school. I went to a Slipknot concert, so a little bit heavier music here in the Bay Area. And I remember at some point the drummer was drumming a cool drum solo and his drum set started raising up. And I was like, whoa, that's cool. And everyone got excited, right? And then halfway through the solo, his drum set started to go like this and tip. And his whole drum set was on the wall while he was upside down playing his drum solo. And everyone was just cheering, like, whoa, this is so crazy, mind-blowing, whoa. It was so impactful, right? I still remember it today. Fast forward, probably about a year ago, I went to New York to see Stranger Things on Broadway, The First Shadow. And it was really cool because the whole time I was there, I was watching it and it felt like it was actually a Netflix episode, but in real life. And the effects that they did there were so, so amazing. I remember even one time near the end, there was a scene where somebody was falling backwards. And it was a slow-motion scene. And it's probably about a minute long, the whole scene. But the fact that they were able to produce that in real life, and it looked like it was a post-production slow-motion scene, was just mind-blowing. So it's like, whoa, that's cool. Given everyone here has probably had some experiences with live performances in the past, I think it's really awesome to see the creativity that people can do with leveraging technology in some creative way. And so I'm thinking about all the projects that we have as engineers over the years where it's like, hey, that's a cool idea. I'm going to write this down. I have a big giant list. Anyone have a list of projects that maybe they'll get to? Yeah, it all happens, right? And it's really awesome because the last six months or so, with AI being able to push us forward with some of these projects, it's just really cool to see where we can take all these projects. So part of my goal today is to inspire you guys to find whatever project you're passionate about and go start building it because it's super easy now with AI. It's easier with AI. I know sometimes there's caveats. But in general, I want to have you guys leave this session inspired to go build whatever that cool thing is for you personally that's going to help you learn stuff or maybe make a difference in the things that you're trying to figure out. So I want to start off with framing this in the realm of the guitar. So the guitar has been around for several years, several thousands of years or hundreds of years, right? And at some point, somebody said, hey, I'm going to put a pickup on a guitar, I'm going to plug it into a speaker, and now we have rock and roll. We're going to make it super loud. That was pretty awesome. And then of course, we have a lot of people getting these stomp boxes, these effects pedals. We put them all together, and we have really awesome sounds. Really famous people have a lot of these. And then at some point, Peter Frampton came along and said, hey, what if we get the sound of the guitar, put it through an actual hose, a physical hose, put it in my mouth, and then play the guitar sound in my mouth and form some sort of words. So that's pretty awesome. That's known as the talk box. And then in the last couple decades, we have had a lot of progress made with software emulation. You think of Pro Tools, Logic, Fruity Loops. There's a lot of software, even to the point where all the effects are in software now. There is the argument, hot take, that maybe you don't need all those physical effects pedals anymore. So there's that question, right? And then of course, looking forward, what is that next evolution of the guitar with AI in the picture? So with that said, I want to take you back several years ago. Around Halloween time, I was passing out candy to a bunch of trick-or-treaters. It got boring. So I was like, hey, what if I go and actually bring my guitar out with my amp and just play? So for the past 10-plus years or so, I've actually been playing guitar on Halloween while passing out candy, and it's kind of fun. And then fast forward probably three years ago, I decided to dress up as Eddie Munson. Anyone know Eddie Munson? Stranger Things guitar player guy. So, pretty awesome guy. So I dressed up as him. I was like, yeah, this is going to be so fun. I could play the music with him, whatever. And you kind of go to—oh, hold on. Hold on, wait a minute. Technical difficulties. All right, there it goes. So, a bunch of little Stranger Things awesomeness, right? But I figured, hey, what if there was more stuff I could put in this whole experience? And so I decided to go and build a little app that would actually paint the Stranger Things alphabet on my garage door, just because it's fun, right? And so let me just show you a quick example of how this worked. So effectively, make sure I have the mic permission. Whenever I play a note, it would go and communicate whatever lights. And I have it to where I could actually set custom messages such as Happy Halloween, all that fun stuff. It was overall pretty fun just to mess with that, right? So once again, whatever I type in there, I could actually spell whatever, and it was very much a cool nod to Stranger Things. But it got me thinking, what is the next evolution of this project? And I settled on this idea of, hey, how hard would it be to make my guitar speak? It sounds easy, maybe, maybe not. So today, I want to share my journey in this process and where I'm at today. So, looking around at the different tools, there's a framework out there called JUCE, really awesome for anyone building audio software. Look into JUCE, it's pretty good. And then of course, there's a number of plugin formats out there for your digital audio workstation. And for those that are not aware, your DAW is effectively your IDE, but for musicians and music producers. And then I started with some text-to-speech stuff with Piper and some built-in Apple stuff, and then some other really fun digital signal processing. So with that said, my first step here was I wanted to get raw text, so I could just type in whatever text I want, push it through the text-to-speech, get the audio clip, and then whenever I play a note on the guitar, I want to go and play that back. So let's go see how that works. So switching over to my Logic Pro here. And this is the plugin I made. So it's just like any other plugin in Logic, where you just pop it in there. It's just chaining all the effects together. And so this is what I came up with. Developers. So it's playing it. Developers. Pretty awesome, right? Developers. Kind of reminds me of something, right? Developers. Start clapping, everyone. Start clapping. Developers. Developers. Developers. Developers. Developers. Developers. Developers. Awesome. Thank you. That's pretty awesome. You guys are great. So I got it to where it's playing an actual audio file. That's pretty awesome. But it turns out, in English, or in any language for that matter, there's more than one word. So the next evolution—oh, that's a little chatty. The next evolution there is, let's actually go and slice it per word now. So I got to the point where it's playing. Now it's going to slice per word. Look at me. I can speak. Look at me. I can speak. So now it's speaking words, and that's pretty awesome, right? But it turns out that as we speak, there are a number of challenges in how we automatically slice words. So I looked into this thing called energy gap segmentation. The general idea here is, if you look at any waveform over here, we see that here's a bunch of words that we're speaking, right? The idea there is that there's typically silence in between words. So let's just cut it whenever the decibels are very much close to zero, right? But the issue with that is there are actually times when, as I'm speaking right now, for example, there's no silence in between some of my words. So it gets a little challenging to where it's not 100% foolproof, right? So beyond that, I looked into this thing called Sonority Peak Syllabifier. That's a hard word to say. Effectively identifying the syllables of the audio signal and identifying that there are vowels in here. Vowels typically lead to syllables. That's kind of the idea. So I said, okay, let's take the sonority peak, add it to the energy gap, and figure out if we could just make that work automatically. So with that said, I made it kind of work. So let's play this one. Thank you for letting me be here. It feels so good to get out of my target case once in a while. So there you go. It's working pretty well. Not quite as good as I want it to. So long story short, I settled on just the ability to go and actually drag this and manually edit some of these segments in here. It worked okay, right? But moving on, in the spirit of evolving the thought, evolving the project, right, it's like, okay, we're having the AI say stuff. That's great. But what if we could actually make it sing? Let's take it to the next step, because this is music. Why not, right? So I looked into pitch detection. And for those that are not aware, as you hear any noise out there, there are typically multiple frequencies going on at any given time. You think of when you play the C key on the piano. There is definitely a fundamental frequency, or the one that we identify as the note, but there's a bunch of other frequencies. So we needed a way to go and figure out, how do we detect that fundamental frequency? So when I press something on the guitar, how do I translate that into an actual note? And then I found this YIN pitch algorithm. Basically, what it does is it detects the pitch. I won't get into all the details, but look it up. It's a really cool way of detecting the pitch. But effectively, what I do is I play the guitar, I detect the pitch, I make what's effectively a pitch sawtooth, or a synthesized note. So for those that are not familiar with how audio works on the computer, you think of all the electronic music, all that stuff. That is basically a synthesized note. We have ADSR, which effectively are the levers to figure out how to actually make the note sound in different ways. And then basically, we get the pitch note, or the synthesized note. We then push it through the voice clip effectively. So think of the talk box. We're kind of filling up the cavity of the voice, that is. And we push it through a vocoder and it should sing. So that's kind of the idea, right? So with that said, let's go ahead and jam out a little bit because I have a guitar and it's fun. So I'm getting back to Logic here. So I'm going to play some chords. I originally was going to play a song that is very much related to the title of my talk, While My Guitar Gently Speaks. But because it's going to be posted online, I don't want to muddy up the waters with any copyright things. So I will play some chords that may or may not sound similar to a famous song. So. So that's the backing track. So let's go ahead and have some AI speaking on top of it. I'm here. I got this. No. Technical difficulties. Let's try that again. Live demos, always the best. I'm here. I got this. All right. Technical difficulties again. So sorry for that. But let's just run through it and see what happens. I'm here. I'm here. I'm here. Friends that speak. I'm a guitar that can speak. Awesome. So there you go. Some bugs to work out. But overall, we are able to say words on top of this. And for what it's worth, I'm trying to mix the synthesized note with this clarity lever right here, so balancing or mixing the synthesized note with the actual voice that the AI is giving. So that worked pretty well. But there's also this other question. Given that when you speak, it's typically conversational, what if I could take this to the next level? And what if we actually had this microphone right here where I could speak into the mic, it would then respond on the guitar? So the way that I accomplished this is, speaking into the mic, use speech to text, so Whisper, put that into raw text, run a local model on my computer. It doesn't really matter what LLM, just any local model, have a conversation coming, and then from that output, go and plop it on the guitar. So let's see how this works. So let's see what it says. Anyone have a question you want to ask my guitar? Anybody? What is reality? Let's think, going through the LLM things, and let's see what it says. That's quite a question to start with. How does your music help you understand that elusive concept of reality right now? So yeah, very existential. I like that. So thank you for that suggestion there. So there you go. We have a choppy version, but it is working. So that's a win, right? Awesome. Thanks for the claps there. But really, it's not quite singing yet, right? So moving on in this project, it's like, how do I actually make it sing, right? I found a lot of open-source options out there as far as recordings or samples, if you will. There's a thing called VocalSet out there. They basically recorded a bunch of singers, and I could use those audio files to go and do some fun stuff with them. And effectively, I took those samples. I used a project called WORLD, which helps with things like pitch shifting and some other things. And I'm able to then shift the pitches of everyone singing on those audio clips and map it to the guitar. This is a very heavy process, and so I can't really do that live. I had to effectively pre-bake that. And then once it's already pre-baked, I could then go and jam out with it. So with that said, let me go ahead and show that example here. So because it takes so long, I actually just started with the vowel sounds. So there's somebody singing each of the five vowels. And so we'll see how that sounds, right? Kind of fun, kind of weird, but overall it's working. It's closer to singing, right? So let's go ahead and throw that on top of all the chords over here. And let's see if we can make it sound decently well. So we'll just play these chords again. So not quite your opera singer, but getting closer. And so I think there's something awesome there. So effectively, once again, that is going through the whole synthesized process, putting a sample, shifting the pitch of the sample, and then effectively mapping each fret or each note on the guitar to one of those samples that are pre-baked effectively, right? So with that said, where do I want to take this next? I very much want to get into more of the AI-heavier options. If anyone has any other ideas of how to make the guitar sing, feel free to come up to me after. I think it's pretty awesome. But more importantly, everyone here, going back to the charge to go and build some awesome side projects, some passion projects of yours, what is that for you? Go and build awesome things. Because nowadays with AI, we could build so many really cool things. And time is typically not the big time suck that it once was, right? So with that said, be awesome, be good to each other, and thank you very much. Thank you very much. I ain't ain't ain't ain't ain't ain't Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye.