I unplugged my Alexa after being woken up at night with Alexa yelling about something or another too many times. It seemed like the perfect noise maker that could also casually perform a limited number of other tasks and parlor tricks when needed but the tech just isnt there.
It has a lot of issues hearing me when I want it to and yet somehow picks up when I'm talking about a friend Alex from a room away. It's not quite as disappointing as the Kindle Fire I bought but at least my girlfriend can use that to watch movies on.
You need to put the payload inside the vehicle before you deliver it. Your chances of successfully installing it later fall asymptotically towards zero.
...which is why your phone keeps getting slower, as it's loaded with code written for better models as a way of marketing their superior capacity. Same reason you can't fit a jet engine to a biplane and expect good performance.
Alexa is a voice interface, and a pretty good one at that. I use it primarily for home automation and it works pretty well for that. There is not much intelligence behind it in that regard, but it's very good at recognizing commands, which is what is needed.
The voice recognition is good, but the interface is completely independent of Alexa. The interface is spawned from a group of people brainstorming all the different ways someone can ask to turn on the lights, and then figuring out the ways to deal with all the different inputs.
I'm not sure how common the knowledge is around making an Alexa skill, but I made one as a side project so I might have some insight.
You define a bunch of variables, and describe what form they are likely to come in (number, data, US city, list of strings). It's not strictly bound to any enumeration you use. Then you list out a bunch of phrases that people are expected to say to activate a function. You can pepper the variables into these phrases. Then you link these phrases to functions, which receive the variables as parameters. What you do from there is up to you.
I would say it's a clever way of gradually reducing the complexity, but it does result in putting the "conversation" complexity on the developer. There is no "state" or compound queries unless the developer thought of the conversation going that way.
Things like Alexa skills existed before Alexa. The magic Amazon added was the hardware + API that was a step ahead of what was there before. Siri by comparison is not as good precisely because the microphones are not as good.
Totally right. I've been building a skill and you quickly learn how un-magical Alexa is. Most of the effort is engineering in fine detail to achieve some degree of natural interaction for your users. For example, adding in state and awareness of what the user said before to influence the skill's behavior and response. It feels like I'm inventing all that from scratch rather than the alexa platform helping achieve that.
If you want to see how well I've done, enable "NextSubway" :)
I've owned the Alexa since their first release. 80% of my request are the weather for the day. Ask for the weather forecast the next few days and she falls short. Other than that she's good for playing music on Spotify or Pandora, and occasionally I'll ask for the time.
I haven't tried Google Home, but the fact that it will read the first search result snippet as an answer seems light years ahead of Alexa.
Side-note: I just made the connection that Amazon bought Alexa.com the web analytics company a long time, and has named Echo's personal assistant Alexa. No coincidence, right? Although they never really mention it.
This article far exaggerates the devices problems, at least as I've experienced it.
If someone shouts my name—Alex—across the apartment, it will activate Alexa
That's trivially fixable. You can rename her to "Echo", "Amazon", or "Computer".
sometimes Alexa will also be activated by arbitrary syllables in ordinary conversation. ... As I read this paragraph aloud to myself, her blue ring has already lit up several times
I've had this happen at random (i.e., when I didn't say her name accidentally) exactly once that I've noticed.
not only is Alexa incapable of looking up even basic facts
This is an exaggeration. By all accounts Google's device does better. But it's simply false that it can't look up "even basic facts". It's quite limited, but far from absolutely useless.
In my mind, Alexa's greatest flaw is in the design for its skill interface. You can't interact with her conversationally, because you always need to preface every command with its skill name, e.g., "Alexa, ask Spaceman what's the news today?" [1]. You must always remember these special names, and used them in this stilted fashion for all interactions. And making this worse is that Amazon insists that the interactions within each skill must be completely discoverable, so that they don't want to publish more than a couple of samples from each, encouraging you to experiment and discover yourself, a process I find tiresome and one guaranteed to leave at least some of the capabilities unknown.
> That's trivially fixable. You can rename her to "Echo", "Amazon", or "Computer".
In my experience, this just changes the times that it accidentally triggers. It doesn't actually reduce how often it happens. Our triggers for random crap all the time, both when we say stuff that sounds like "Alexa" and for completely indecipherable reasons. We actually have two of them in our house, and only one regularly goes of at the wrong time. I thought it might be defective, so I swapped them, but then the other one started doing it. It's location-dependent in our house, so it must have to do with the acoustics or something. Regardless, it's annoying, and the less technical family members are far less tolerant of this kind of stuff than I am.
All that said, I still agree that the article goes overboard in exaggeration. Overall I like my Echos, and I think they work pretty well most of the time. There's a lot of room for improvement, but it is a new product.
While mine has never 'accidentally' ordered anything or yelled about internet disconnect, I do agree with the conclusion that it ends up being a very specific device. For me, that is a connected speaker with good sound that I can gently manage from across the room. Good enough for me.
On another note, my 2 year old son has always had a complicated relationship with 'her'. At one point, she red-ringed and would issue inscrutable messages after long periods of spinning white. This really sent him into a tailspin and it has taken several months for him to trust her again.
This is fascinating. It hadn't dawned on me that very small children growing up with these in their homes will have a very human response to them. A relationship regardless of the actual intelligence of the device.
Yet you can't name the timers, so it's easy to forget which is which when you set timers staggered throughout cooking at different intervals (even more exacerbated if you dare try to time things to generally come together at the same time).
Ironically this has been +1'd so many time on Amazon's own forums... starting years ago.
I spent a minute staring at this comment in puzzlement wondering what party to the interaction this "Edit" person is and what hard drive they were formatting while telling the world about it.
I have an Amazon Echo in every room of my house, and I use it to control home automation, such as all of my lights, my ecobee thermostat, and my entire home theater (via a logitech harmony). I even use it to switch on and off my humidifier and space heater (and manage the temperature). My wife and I also use it to manage my shopping list (via the OurGroceries skill/app). I can use it to remote start my car and preheat or cool it down in the morning. I can even control my adjustable tempurpedic bed base, and have an IFTTT trigger which makes my bed vibrate whenever my amazon echo alarm goes off in the morning.
The possibilities are endless, and it's only going to get better.
I have my home automation stuff hooked up to Alexa also, and no, it's not always easier to do it with my hands. Sometimes it is, sometimes it isn't. I am often baffled by this "I'll just use the light switch" response to home automation. It's not that you can't "just use the light switch," but there are situations where it makes life measurably better to not have to. First world problems, for sure, but they exist.
For example, if I'm sitting comfortably in my recliner and I decide I want to watch a movie, it's pretty great to be able to tap a few buttons on my phone to dim the lights without getting up. It's even better to now be able to say "alexa, turn on movie lights" to do it without digging out and unlocking my phone.
Or if I'm making dinner, feeling hot, and my hands are covered in tomato juice, it's awfully nice to be able to adjust the air conditioner by voice.
These aren't life changing things, but they are pretty nice luxuries that have increased my family's quality of life.
I don't get the fascination with this tech. The computer in Star Trek at least had a personality and intonation in its voice when it responded, even though it sounded like a synthesized voice.
The key here is intonation -- the infinitely varying way in which we vary the speed and pitch of our voices in response to another human being. It's very subtle and something Amaazon, Google and Siri have been unable to capture. And it makes interacting with these interfaces awkward and annoying to me. Either it needs to be perfect or it gets relegated to a "that's neat" category and promptly shut off.
When I talk to one of these things, it's downright painful. It's like speaking to a someone who's (at least interacting in english, I'm sure its true of other languages) primary language isn't english and they are just saying a series of words without any meaning behind them. You tolerate this (although its still annoying) because you're interacting with a human being on the other end, not some white or beige box.
The issue behind this of course is these devices will likely never be able to master intonation, because doing so would be a feat that would require clear understanding of subtle context, far greater than what current NLP technology can do, and may even require a general intelligence. It might be possible, I'm not up to date on the current state of the art but certainly wouldn't be computationally feasible in a consumer product any time soon.
Until then, I'll never buy or use one of these stupid voice interfaces. I can get the same task done on my phone or whatever UI without the annoyance.
I understand what you mean, but for me, she has developed something of a personality. It is kind of a bumbling but well intentioned friend. Similar to my Roomba which is not the best at what he does, but keeps on trying (i.e. bumping futilely into the same walls).
Google Home is starting to get there, especially in the games where the total set of commands and responses is much lower. It definitely has room to grow, but it'll get there given time. VR was terrible in the 90s, now it's breathtaking.
I think the value of these systems is when you have your hands full, or when getting out your phone or other UI is not convenient.
"Alexa, turn on the lights".
"Alexa, let the cat out".
"Alexa, did I lock the door?"
As soon as there are enough of these, it'll take off because it'll be useful. None of these things need real intelligence. Sure, intonation won't be great. But the convenience will override the imperfection.
That's all I would ever want from an AI. I don't even need complicated answers, just a repeat of the requested task and maybe a confirmation before going ahead. More like air traffic control than a nuanced buttler.
That also lowers the possibility of it developing emotions and trying to subvert global politics from my toaster.
The issue of intonation, here, isn't a technical one. I did my third years' internship in a 3 people startup, lost in the middle of the french countryside. And we were working on a software that could generate complete weather bulletin for any european city (>50k inhabitants) or simply deliver a personalized voicemail (with names, dates, past purchases) like any human would do. In fact people were calling back the number we gave, asking for "jeremy", the bot. It was really good and it was less than 1k lines of python code (embedded in django).
So, if Amazon wanted it to be something more than a marketing fluff then could -probably- do something about it.
That's actually pretty neat! So how come other voices don't match this? Is it the fact that intonation in a single domain (weather) is easier to predict than if the voice needs to handle general (broad) topics of conversation?
It's not about the topic, but mostly the technique. You have to record and process a few hours of human speech in a studio (when I was there the "studio" was a closet with blankets on the wall). Most of the words are not cut in the same way as others text-to-speech engine would do, to preserve intonation, hence longer recording times and a "grammar" limited to the the span of the recorded vocabulary (see if you can catch how the topic is always about the rain. Maybe it was due to us being in britany, a very rainy french region)
However, given the number of complaints here, I don't think a cap on the number of topics she can discuss would be an issue for Alexa.
Is that the same as Talking Sam on the C64? That was the first synthesized speech I ever got to play with.
I also had the speech module on my TI-99/4A. That was pretty sweet. That led to me buying a bunch of SPO256 chips from Radio Shack that I eventually wired to up to a little line-following robot that I made. I was the only one that could understand it.
I don't get this rant. Alexa is still bleeding edge. It does some things really well. It does some things really poorly. There are lots of things it can't do. So the main complaint is that Alexa could be better, your point..?
To me, right now it excels as an interface when you are busy doing something else. Ours is used most in the kitchen and living room. Timers, streaming music/audio and basic facts all work great. Home automation is improving and will get there as systems learn to play nice. That seems like a lot for $50, but maybe that's just me...
The killer app hasn't been created yet (and possibly not even envisioned). I can't wait until they are connected and I can issue a command to move streaming from one device to another one in another room ("Alexa, pause for Jim", "Alexa, continue for Jim").
I think the biggest downside of voice UI right now is a lack of authentication. Someone outside the window could yell "Alexa, open the door".
> I can't wait until they are connected and I can issue a command to move streaming from one device to another one in another room
Google Home can do this to a degree today. You are capable of saying things like "Play X on my TV." if your TV is wired up with a ChromeCast.
In general, Google Home is very far ahead of Alexa. I agree with the opinion that Alexa is a glorified radio, but I don't believe that opinion extends to all devices in that class. I think its just because Alexa is overrated and not very powerful.
You give a really vanilla use case and than suggest Google Home is superior. Can you give me a use case where it actually demonstrates it is far ahead.
The Chromecast example isn't supporting your claim in my opinion.
Entertainment is the biggest use care for these devices IMO; Music and TV. I haven't used an Echo in a while, but it had nothing for TV and wasn't that great at music at the time. Does it have good support for both now?
I always try all these things as I am a junky like that but these things seem to have such limited use cases. Looking up and playing things on chromecast, for me, is faster using my phone than speech will ever be. So is controlling things in the house and with a phone I can open the catflap without waking my wife talking to Alexa. Switching on and off things is okish (phone or even a normal lightswitch are still faster and less hassle imho), but trying to look up what that great song from that obscure Norwegian black metal band was, of which the name is something Alexa does not parse and will take me 10 minutes to properly explain while it makes me sound like robot doing it; while via my phone I google it in seconds. I am really not convinced if this is not just gadgets because of gadgets. Or maybe I am just weird. Which is fine too.
Where it really shines for me is playing the latest podcast on YouTube or continuing a Netflix show where I left off.
Being able to say "Watch Breaking bad on the living room TV" and have Netflix pop-up and start playing right where I left off, or being able to say "play the latest off topic podcast on the bedroom TV" and have it startup right away is really nice.
I've also got a lot of home automation stuff so that combined with telling it to turn lights on/off is fantastic.
My nightly routine is to tell it to turn the bedroom lights on, get ready for bed, tell it to play the show I'm watching, then when I'm in bed and comfy tell it to turn all lights off.
I don't think it's fully there yet, but I really disagree that it's more convenient to use a phone app than a Google Home. Alexa is worse here since skills are clunky, and while I agree that voice isn't the best setup in bedrooms with multiple people, it's definitely easier to say "Hey Google, open catflap" than futzing with a phone.
For me, simply being able to turn my TV + soundbar on/off with Google Home + IFTTT + Harmony Hub, while being a stupid amount of duct tape, has made me no longer need to hunt for where the TV or sound bar remote are any more. Being able to pause/resume shows on my Roku with the same setup avoids me having to get my phone out, unlock it, open the Roku app and press the pause button. It's pretty convenient. Is it life changing? No. Could I just stop losing the remotes? No :p
I have also found myself listening to a lot more music with the google home since I can just turn it on with my voice when I crawl out of bed. Not sure if that's a good fit for metal of any kind though.
I don't think it's truly life changing, few things are, but well executed voice control is truly convenient.
Keyword being yet. Recently Plex announced integration with Alexa. Nvidia Shield recently saw Amazon Video app. It is really just a matter of time before skills to play video on devices like the audio is available.
> I think the biggest downside of voice UI right now is a lack of authentication. Someone outside the window could yell "Alexa, open the door".
I think this is the most overlooked use case that needs to be addressed. Even in a secure residence, if you leave a 2nd story window open it might be possible to get the audio to Alexa. Though I suppose these skills could add a 2-factor to the unlocking mechanism.
"Alexa unlock the door."
"Ok, what is the passcode."
<get's out phone, retrieves token>
"123456"
"Thanks, unlocking door."
The problem I see is I don't want 2-factor on everything.
Now that I think of it, as someone that has 3 different 2f tokens.. smart watch integration would actually be useful in general... If I had a smart watch.
I'd like to be personally involved in some transactions. I think this will be a major use case for smart watches and a major feature of home interfaces (phone, smartwatch, or whatever has the totp) in the future.
A smartwatch as simple as a Pebble can run such an app already. For example, I authenticate with 2FA with e.g. Google & Facebook w/my Pebble.
The problem is that I don't want to use my voice for commanding that much. It has privacy implications. First, obviously there is data being transmitted to a US corporation such as Google or Amazon. Second, there is analog audio data being transmitted. For example, that TOTP token generally stays valid for 30 seconds. In this window, an attacker could open the door right after you entered. I live in an apartment, and anything I say in the hallway can be heard by other people living on the same level as me while they are within their house (yep, terrible design).
I do admit that in the privacy of your home such analog audio data works better than say in public via your smartphone. But then again, smartphone is a can of worms with more privacy implications such as screen easily being seen by bystanders and cameras (a smartwatch has the very same issue; Google Glass wouldn't!). For the phone, Google Home is an interface, and certainly not (yet) the primary interface. For Amazon Echo, Alexa is the interface.
Bottomline, we need to think of these disadvantages (share our thoughts about them), and understand when they're too big to be a valid use case.
I mean people can already bump most locks, or break a window, or in your case just climb in the window...
No house is really secure, and I'm not going to spend my time trying to secure it against non-existent threats.
That's not to say I don't want security. The threat of someone finding a loophole and being able to enumerate and unlock all smart doors is very real. But I'm not going to care if someone can come up with a scheme where they megaphone their voice to unlock the door, then somehow prevent the notification going to my phone letting me know someone is home, avoiding or destroying the few security cameras we have, and avoiding being caught by someone else during all this ruckus. Especially when they could just break a window...
> I can't wait until they are connected and I can issue a command to move streaming from one device to another one in another room ("Alexa, pause for Jim", "Alexa, continue for Jim").
Too much overhead. It should be able to identify you by voice, so you can just ask it to "pause" and "resume".
While I think this is necessary, I don't think it's sufficient. Sure, when your smart home device is just hooked up to some lights and maybe your thermostat, it's fine. But pretty soon (and/or possibly now, I haven't kept up with the news), your smart home device will be connected to things that warrant a little more security, e.g. door locks. It's not that hard for someone to get a recording of your voice.
I don't get this rant. Alexa is still bleeding edge.
But it's being sold to relatively unsophisticated consumers, on the front page of the world's largest e-commerce site. This is not a bleeding edge release like Google Glass; this is a broad market release to all kinds of users.
Like I said in my other comment, I do like mine, and I think the article goes overboard. But I don't think Amazon has the luxury of saying "hey, it's bleeding edge!" when it's a top selling item for them, and front and center when you go to their site.
How about if it already knows you're Jim? I think understanding context is holding it back. Azure Cognitive Services has a speaker recognition service. I'm hopeful that Amazon will get there.
I only got the Dot a couple of weeks ago and for me it's worth the $50 just for playing music. Anyone who thinks it is a "glorified radio clock" might want to just take some time to think about exactly what it's doing. I'm kind of amazed by it.
The title is making the assumption that the future of AI isn't glorified radio clock.
Also, what's definition of "future". Academic research paper or project? Or, something deployed and used by millions. I tend towards the later. And by that metric the future is now.
I think they really need to work on how to report problems succinctly. A simple error tone would be fine, combined with some recognized follow-up command like “what happened” or “tell me more”; in other words, assume by default that the person probably just mumbled and is about to repeat the real command.
Instead, Alexa likes to spew out multi-sentence explanations when there is a problem, and trying to talk over all of that to get Alexa’s attention again is very irritating. Usually I end up trying to talk at the same time and it may or may not listen. Imagine Alexa saying all of this kind of stuff: “I am sorry, the Thing I Think You Asked Me About But You Did Not Actually Ask Me About is not available right now. I cannot find a solution. Please go to Amazon dot com slash Something Else You Most Definitely Did Not Ask About, and visit the solutions web page.”. While all of that unnecessary verbiage is being barked out, I am usually saying something like: “Alexa...Alexa...ALEXA, turn on the...ALEXA, turn on the...ALEXA STOP!!! Alexa, turn on the lights.”.
This guy seems to be most upset about not being able to search Google verbally. Alexa has its skills, but random search is not its finest display of talent (Bing?)
This article misses the point. Alexa is interesting because it's a new user interface (voice + always listening), not because it's "an AI". Imo, it's the biggest UI jump since the touchscreen phone. The tech that enables the UI will continue to improve, but it doesn't need to be a general purpose conversational AI to be successful.
Well, a good point that there is no AI except for speech recognition, but, overlooks the utility of hands-free operation. In scenarios like cooking, in jacuzzi/bath/shower, or just when you are some distance from the device, it is useful to control listening to audio books, music, NPR radio, etc., my family owns two Echoes and get good use from them.
I was at a multifamily technology conference in San Diego last year and during a panel with some 'experts', brought up the concern of devices like Amazon Echo recording people's conversations. The lady who responded misunderstood me and chimed in within something like: "you know, I haven't even thought of that. There are many opportunities for revenue with this technology". But that was an evil uber-landlord, not sure where Amazon stands on this. If they want to increase adoption for this technology, they better come up with a solid privacy policy and stand behind it.
And technologically the language processing "necessarily" takes place in the cloud. For that everything you say needs to be sent through the web. Amazon will store everything they can for the purpose of later using this as a training data set.
Have you considered that maybe other people have thought about this and reached a different conclusion? I think calling it an NSA bug is pretty silly if you have a smartphone or laptop. How are those not bugs? All of these devices have microphones.
The Echo is less likely to be used as a bug than any other device because it's simple and single-purpose. When you are not using it, it is idle and will not use network throughput. That's trivially easy to monitor with my router. I can see how many gigabytes of data my TV downloads. If an Echo started consuming network throughput to upload audio when I was not using it, then that would be extremely suspicious and easy to notice. This ease of detection makes it a terrible covert surveillance device.
By comparison, I have little visibility into what my phone is streaming over its cellular connection, and devices like laptops, PCs, and Xboxes are so complicated that I cannot expect their network usage to be idle. It is unsurprising for these devices to download/upload potentially gigabytes of data while not in use, because some software has decided to patch in the background. These devices have local storage and could buffer audio and upload it during times of legitimate network activity. There is a lot of network activity in which covert surveillance could hide. This is not so with the Echo.
Additionally, your phone and laptop are far more juicy targets for the NSA or any attacker. These devices have your data! An attacker can break in and steal your documents or passwords or other data. And then they can also activate the microphone and listen to you! The Echo doesn't have any data on it, so it's not a juicy target. It's also stationary in one room and has far less ability to conduct effective surveillance on someone than the cell phone or laptop that they carry with them everywhere. Someone isn't going to take their Echo with them to the mosque or an important business meeting or whatever.
The Echo is far less complicated than smartphones and laptops (less stuff running on it), and so is less likely to have software defects that can be exploited remotely by attackers.
With phone, the way it currently works (Apple, Google essentially have full control over it, the modem cpu also is closed source and has full control over application cpu) it's not hard and people are concerned about it.
With laptop (at least non mac) is a bit harder, different hardware, different OSes. Some laptops have privacy switches. Regarding turning speakers into a microphone this is possible, but turning D/A to A/D converter is much harder.
I noticed that Alexa was talking a lot to Amazon servers even when no one was at home (i.e. no one was even using it), which made me suspicious and I just unplugged it. Anyway I wasn't using it lately for anything else than an alarm clock, the novelty worn out after a week or so.
This article is a case of over expectations. Echo IS a great clock radio and that's fine with me. The great part about it is the ability to access my music. If I want to hear x all I have to do is ask for it. 5 years ago I would never have had expected to be able to that for less than $50.00. Everything else that comes with it is just icing on the cake. We are still years away, if ever, from the Star Trek computer. And that's fine. Keep your expectations in check and you can see that it's excellent technology --with a great future.
84 comments
[ 2.8 ms ] story [ 149 ms ] threadIt has a lot of issues hearing me when I want it to and yet somehow picks up when I'm talking about a friend Alex from a room away. It's not quite as disappointing as the Kindle Fire I bought but at least my girlfriend can use that to watch movies on.
"Why isn't your desk in front of the telescreen?"
Arrrggghhh! The problem of course was that the Echo rebooted faster than the wireless router after a power failure.
Is there really no setting on this thing to make it shut down from 11PM to 6AM?
I'm not sure how common the knowledge is around making an Alexa skill, but I made one as a side project so I might have some insight.
You define a bunch of variables, and describe what form they are likely to come in (number, data, US city, list of strings). It's not strictly bound to any enumeration you use. Then you list out a bunch of phrases that people are expected to say to activate a function. You can pepper the variables into these phrases. Then you link these phrases to functions, which receive the variables as parameters. What you do from there is up to you.
I would say it's a clever way of gradually reducing the complexity, but it does result in putting the "conversation" complexity on the developer. There is no "state" or compound queries unless the developer thought of the conversation going that way.
If you want to see how well I've done, enable "NextSubway" :)
https://www.amazon.com/Matt-Sahn-NextSubway/dp/B01N9MO4DT/
> in order to use the Hue bulbs with Alexa, you must relinquish all use of manual light switches.
Isn't fully true. We have smart dimmers that can be placed in better spots since they are battery powered.
I haven't tried Google Home, but the fact that it will read the first search result snippet as an answer seems light years ahead of Alexa.
Side-note: I just made the connection that Amazon bought Alexa.com the web analytics company a long time, and has named Echo's personal assistant Alexa. No coincidence, right? Although they never really mention it.
If someone shouts my name—Alex—across the apartment, it will activate Alexa
That's trivially fixable. You can rename her to "Echo", "Amazon", or "Computer".
sometimes Alexa will also be activated by arbitrary syllables in ordinary conversation. ... As I read this paragraph aloud to myself, her blue ring has already lit up several times
I've had this happen at random (i.e., when I didn't say her name accidentally) exactly once that I've noticed.
not only is Alexa incapable of looking up even basic facts
This is an exaggeration. By all accounts Google's device does better. But it's simply false that it can't look up "even basic facts". It's quite limited, but far from absolutely useless.
In my mind, Alexa's greatest flaw is in the design for its skill interface. You can't interact with her conversationally, because you always need to preface every command with its skill name, e.g., "Alexa, ask Spaceman what's the news today?" [1]. You must always remember these special names, and used them in this stilted fashion for all interactions. And making this worse is that Amazon insists that the interactions within each skill must be completely discoverable, so that they don't want to publish more than a couple of samples from each, encouraging you to experiment and discover yourself, a process I find tiresome and one guaranteed to leave at least some of the capabilities unknown.
[1] Shameless plug - this is a skill I wrote to give you news about scheduled rocket launches, near earth object passes, and lunar phase - https://sites.google.com/view/spacemanforalexa/
In my experience, this just changes the times that it accidentally triggers. It doesn't actually reduce how often it happens. Our triggers for random crap all the time, both when we say stuff that sounds like "Alexa" and for completely indecipherable reasons. We actually have two of them in our house, and only one regularly goes of at the wrong time. I thought it might be defective, so I swapped them, but then the other one started doing it. It's location-dependent in our house, so it must have to do with the acoustics or something. Regardless, it's annoying, and the less technical family members are far less tolerant of this kind of stuff than I am.
All that said, I still agree that the article goes overboard in exaggeration. Overall I like my Echos, and I think they work pretty well most of the time. There's a lot of room for improvement, but it is a new product.
On another note, my 2 year old son has always had a complicated relationship with 'her'. At one point, she red-ringed and would issue inscrutable messages after long periods of spinning white. This really sent him into a tailspin and it has taken several months for him to trust her again.
Ironically this has been +1'd so many time on Amazon's own forums... starting years ago.
Me: Alexa, set up an alarm for 2 minutes
Alexa: Two minutes.
[Two minutes later]
Alexa: Beep beep beep
Me: Alexa, thank you!
Alexa: You are welcome... Beep beep beep...
Edit: formatting
I have come to realize I mumble and talk low and cold heartless Alexa will remind me to speak up.
That being said, I never thought of my Echo as anything other than a 1st step towards something greater.
The possibilities are endless, and it's only going to get better.
For example, if I'm sitting comfortably in my recliner and I decide I want to watch a movie, it's pretty great to be able to tap a few buttons on my phone to dim the lights without getting up. It's even better to now be able to say "alexa, turn on movie lights" to do it without digging out and unlocking my phone.
Or if I'm making dinner, feeling hot, and my hands are covered in tomato juice, it's awfully nice to be able to adjust the air conditioner by voice.
These aren't life changing things, but they are pretty nice luxuries that have increased my family's quality of life.
Seriously, man? I get that you want to justify the thousands you spent on home automation, but come on.
The key here is intonation -- the infinitely varying way in which we vary the speed and pitch of our voices in response to another human being. It's very subtle and something Amaazon, Google and Siri have been unable to capture. And it makes interacting with these interfaces awkward and annoying to me. Either it needs to be perfect or it gets relegated to a "that's neat" category and promptly shut off.
When I talk to one of these things, it's downright painful. It's like speaking to a someone who's (at least interacting in english, I'm sure its true of other languages) primary language isn't english and they are just saying a series of words without any meaning behind them. You tolerate this (although its still annoying) because you're interacting with a human being on the other end, not some white or beige box.
The issue behind this of course is these devices will likely never be able to master intonation, because doing so would be a feat that would require clear understanding of subtle context, far greater than what current NLP technology can do, and may even require a general intelligence. It might be possible, I'm not up to date on the current state of the art but certainly wouldn't be computationally feasible in a consumer product any time soon.
Until then, I'll never buy or use one of these stupid voice interfaces. I can get the same task done on my phone or whatever UI without the annoyance.
"Alexa, turn on the lights".
"Alexa, let the cat out".
"Alexa, did I lock the door?"
As soon as there are enough of these, it'll take off because it'll be useful. None of these things need real intelligence. Sure, intonation won't be great. But the convenience will override the imperfection.
That also lowers the possibility of it developing emotions and trying to subvert global politics from my toaster.
So, if Amazon wanted it to be something more than a marketing fluff then could -probably- do something about it.
Is it good enough for you? :)
Note: I worked there in 2013, they are smart people. And yes, these are _computer generated_ messages, with human voices.
However, given the number of complaints here, I don't think a cap on the number of topics she can discuss would be an issue for Alexa.
I also had the speech module on my TI-99/4A. That was pretty sweet. That led to me buying a bunch of SPO256 chips from Radio Shack that I eventually wired to up to a little line-following robot that I made. I was the only one that could understand it.
To me, right now it excels as an interface when you are busy doing something else. Ours is used most in the kitchen and living room. Timers, streaming music/audio and basic facts all work great. Home automation is improving and will get there as systems learn to play nice. That seems like a lot for $50, but maybe that's just me...
The killer app hasn't been created yet (and possibly not even envisioned). I can't wait until they are connected and I can issue a command to move streaming from one device to another one in another room ("Alexa, pause for Jim", "Alexa, continue for Jim").
I think the biggest downside of voice UI right now is a lack of authentication. Someone outside the window could yell "Alexa, open the door".
Google Home can do this to a degree today. You are capable of saying things like "Play X on my TV." if your TV is wired up with a ChromeCast.
In general, Google Home is very far ahead of Alexa. I agree with the opinion that Alexa is a glorified radio, but I don't believe that opinion extends to all devices in that class. I think its just because Alexa is overrated and not very powerful.
The Chromecast example isn't supporting your claim in my opinion.
Being able to say "Watch Breaking bad on the living room TV" and have Netflix pop-up and start playing right where I left off, or being able to say "play the latest off topic podcast on the bedroom TV" and have it startup right away is really nice.
I've also got a lot of home automation stuff so that combined with telling it to turn lights on/off is fantastic.
My nightly routine is to tell it to turn the bedroom lights on, get ready for bed, tell it to play the show I'm watching, then when I'm in bed and comfy tell it to turn all lights off.
For me, simply being able to turn my TV + soundbar on/off with Google Home + IFTTT + Harmony Hub, while being a stupid amount of duct tape, has made me no longer need to hunt for where the TV or sound bar remote are any more. Being able to pause/resume shows on my Roku with the same setup avoids me having to get my phone out, unlock it, open the Roku app and press the pause button. It's pretty convenient. Is it life changing? No. Could I just stop losing the remotes? No :p
I have also found myself listening to a lot more music with the google home since I can just turn it on with my voice when I crawl out of bed. Not sure if that's a good fit for metal of any kind though.
I don't think it's truly life changing, few things are, but well executed voice control is truly convenient.
I think this is the most overlooked use case that needs to be addressed. Even in a secure residence, if you leave a 2nd story window open it might be possible to get the audio to Alexa. Though I suppose these skills could add a 2-factor to the unlocking mechanism.
"Alexa unlock the door."
"Ok, what is the passcode."
<get's out phone, retrieves token>
"123456"
"Thanks, unlocking door."
The problem I see is I don't want 2-factor on everything.
"Alexa unlock the door."
"Ok, what is the passcode."
<Looks at smart watch>
"123456"
"Thanks, unlocking door."
Now that I think of it, as someone that has 3 different 2f tokens.. smart watch integration would actually be useful in general... If I had a smart watch.
The problem is that I don't want to use my voice for commanding that much. It has privacy implications. First, obviously there is data being transmitted to a US corporation such as Google or Amazon. Second, there is analog audio data being transmitted. For example, that TOTP token generally stays valid for 30 seconds. In this window, an attacker could open the door right after you entered. I live in an apartment, and anything I say in the hallway can be heard by other people living on the same level as me while they are within their house (yep, terrible design).
I do admit that in the privacy of your home such analog audio data works better than say in public via your smartphone. But then again, smartphone is a can of worms with more privacy implications such as screen easily being seen by bystanders and cameras (a smartwatch has the very same issue; Google Glass wouldn't!). For the phone, Google Home is an interface, and certainly not (yet) the primary interface. For Amazon Echo, Alexa is the interface.
Bottomline, we need to think of these disadvantages (share our thoughts about them), and understand when they're too big to be a valid use case.
No house is really secure, and I'm not going to spend my time trying to secure it against non-existent threats.
That's not to say I don't want security. The threat of someone finding a loophole and being able to enumerate and unlock all smart doors is very real. But I'm not going to care if someone can come up with a scheme where they megaphone their voice to unlock the door, then somehow prevent the notification going to my phone letting me know someone is home, avoiding or destroying the few security cameras we have, and avoiding being caught by someone else during all this ruckus. Especially when they could just break a window...
Too much overhead. It should be able to identify you by voice, so you can just ask it to "pause" and "resume".
But it's being sold to relatively unsophisticated consumers, on the front page of the world's largest e-commerce site. This is not a bleeding edge release like Google Glass; this is a broad market release to all kinds of users.
Like I said in my other comment, I do like mine, and I think the article goes overboard. But I don't think Amazon has the luxury of saying "hey, it's bleeding edge!" when it's a top selling item for them, and front and center when you go to their site.
Also, what's definition of "future". Academic research paper or project? Or, something deployed and used by millions. I tend towards the later. And by that metric the future is now.
Instead, Alexa likes to spew out multi-sentence explanations when there is a problem, and trying to talk over all of that to get Alexa’s attention again is very irritating. Usually I end up trying to talk at the same time and it may or may not listen. Imagine Alexa saying all of this kind of stuff: “I am sorry, the Thing I Think You Asked Me About But You Did Not Actually Ask Me About is not available right now. I cannot find a solution. Please go to Amazon dot com slash Something Else You Most Definitely Did Not Ask About, and visit the solutions web page.”. While all of that unnecessary verbiage is being barked out, I am usually saying something like: “Alexa...Alexa...ALEXA, turn on the...ALEXA, turn on the...ALEXA STOP!!! Alexa, turn on the lights.”.
Are you people aware that this is essentially a glorified bug?
And technologically the language processing "necessarily" takes place in the cloud. For that everything you say needs to be sent through the web. Amazon will store everything they can for the purpose of later using this as a training data set.
Is that actually true? And if so, how far are we from being able to do this on the device itself?
The Echo is less likely to be used as a bug than any other device because it's simple and single-purpose. When you are not using it, it is idle and will not use network throughput. That's trivially easy to monitor with my router. I can see how many gigabytes of data my TV downloads. If an Echo started consuming network throughput to upload audio when I was not using it, then that would be extremely suspicious and easy to notice. This ease of detection makes it a terrible covert surveillance device.
By comparison, I have little visibility into what my phone is streaming over its cellular connection, and devices like laptops, PCs, and Xboxes are so complicated that I cannot expect their network usage to be idle. It is unsurprising for these devices to download/upload potentially gigabytes of data while not in use, because some software has decided to patch in the background. These devices have local storage and could buffer audio and upload it during times of legitimate network activity. There is a lot of network activity in which covert surveillance could hide. This is not so with the Echo.
Additionally, your phone and laptop are far more juicy targets for the NSA or any attacker. These devices have your data! An attacker can break in and steal your documents or passwords or other data. And then they can also activate the microphone and listen to you! The Echo doesn't have any data on it, so it's not a juicy target. It's also stationary in one room and has far less ability to conduct effective surveillance on someone than the cell phone or laptop that they carry with them everywhere. Someone isn't going to take their Echo with them to the mosque or an important business meeting or whatever.
The Echo is far less complicated than smartphones and laptops (less stuff running on it), and so is less likely to have software defects that can be exploited remotely by attackers.
Maybe nobody has brought this up because it's complete nonsense.
With laptop (at least non mac) is a bit harder, different hardware, different OSes. Some laptops have privacy switches. Regarding turning speakers into a microphone this is possible, but turning D/A to A/D converter is much harder.