Skip to main content

Blog

Learn About Our Meetup

5000+ Members

MEETUPS

LEARN, CONNECT, SHARE

Join our meetup, learn, connect, share, and get to know your Toronto AI community. 

JOB POSTINGS

INDEED POSTINGS

Browse through the latest deep learning, ai, machine learning postings from Indeed for the GTA.

CONTACT

CONNECT WITH US

Are you looking to sponsor space, be a speaker, or volunteer, feel free to give us a shout.

Author: torontoai

An All-Neural On-Device Speech Recognizer

In 2012, speech recognition research showed significant accuracy improvements with deep learning, leading to early adoption in products such as Google’s Voice Search. It was the beginning of a revolution in the field: each year, new architectures were developed that further increased quality, from deep neural networks (DNNs) to recurrent neural networks (RNNs), long short-term memory networks (LSTMs), convolutional networks (CNNs), and more. During this time, latency remained a prime focus — an automated assistant feels a lot more helpful when it responds quickly to requests.

Today, we’re happy to announce the rollout of an end-to-end, all-neural, on-device speech recognizer to power speech input in Gboard. In our recent paper, “Streaming End-to-End Speech Recognition for Mobile Devices“, we present a model trained using RNN transducer (RNN-T) technology that is compact enough to reside on a phone. This means no more network latency or spottiness — the new recognizer is always available, even when you are offline. The model works at the character level, so that as you speak, it outputs words character-by-character, just as if someone was typing out what you say in real-time, and exactly as you’d expect from a keyboard dictation system.

This video compares the production, server-side speech recognizer (left panel) to the new on-device recognizer (right panel) when recognizing the same spoken sentence. Video credit: Akshay Kannan and Elnaz Sarbar

A Bit of History
Traditionally, speech recognition systems consisted of several components – an acoustic model that maps segments of audio (typically 10 millisecond frames) to phonemes, a pronunciation model that connects phonemes together to form words, and a language model that expresses the likelihood of given phrases. In early systems, these components remained independently-optimized.

Around 2014, researchers began to focus on training a single neural network to directly map an input audio waveform to an output sentence. This sequence-to-sequence approach to learning a model by generating a sequence of words or graphemes given a sequence of audio features led to the development of “attention-based” and “listen-attend-spell” models. While these models showed great promise in terms of accuracy, they typically work by reviewing the entire input sequence, and do not allow streaming outputs as the input comes in, a necessary feature for real-time voice transcription.

Meanwhile, an independent technique called connectionist temporal classification (CTC) had helped halve the latency of the production recognizer at that time. This proved to be an important step in creating the RNN-T architecture adopted in this latest release, which can be seen as a generalization of CTC.

Recurrent Neural Network Transducers
RNN-Ts are a form of sequence-to-sequence models that do not employ attention mechanisms. Unlike most sequence-to-sequence models, which typically need to process the entire input sequence (the waveform in our case) to produce an output (the sentence), the RNN-T continuously processes input samples and streams output symbols, a property that is welcome for speech dictation. In our implementation, the output symbols are the characters of the alphabet. The RNN-T recognizer outputs characters one-by-one, as you speak, with white spaces where appropriate. It does this with a feedback loop that feeds symbols predicted by the model back into it to predict the next symbols, as described in the figure below.

Representation of an RNN-T, with the input audio samples, x, and the predicted symbols y. The predicted symbols (outputs of the Softmax layer) are fed back into the model through the Prediction network, as yu-1, ensuring that the predictions are conditioned both on the audio samples so far and on past outputs. The Prediction and Encoder Networks are LSTM RNNs, the Joint model is a feedforward network (paper). The Prediction Network comprises 2 layers of 2048 units, with a 640-dimensional projection layer. The Encoder Network comprises 8 such layers. Image credit: Chris Thornton

Training such models efficiently was already difficult, but with our development of a new training technique that further reduced the word error rate by 5%, it became even more computationally intensive. To deal with this, we developed a parallel implementation so the RNN-T loss function could run efficiently in large batches on Google’s high-performance Cloud TPU v2 hardware. This yielded an approximate 3x speedup in training.

Offline Recognition
In a traditional speech recognition engine, the acoustic, pronunciation, and language models we described above are “composed” together into a large search graph whose edges are labeled with the speech units and their probabilities. When a speech waveform is presented to the recognizer, a “decoder” searches this graph for the path of highest likelihood, given the input signal, and reads out the word sequence that path takes. Typically, the decoder assumes a Finite State Transducer (FST) representation of the underlying models. Yet, despite sophisticated decoding techniques, the search graph remains quite large, almost 2GB for our production models. Since this is not something that could be hosted easily on a mobile phone, this method requires online connectivity to work properly.

To improve the usefulness of speech recognition, we sought to avoid the latency and inherent unreliability of communication networks by hosting the new models directly on device. As such, our end-to-end approach does not need a search over a large decoder graph. Instead, decoding consists of a beam search through a single neural network. The RNN-T we trained offers the same accuracy as the traditional server-based models but is only 450MB, essentially making a smarter use of parameters and packing information more densely. However, even on today’s smartphones, 450MB is a lot, and propagating signals through such a large network can be slow.

We further reduced the model size by using the parameter quantization and hybrid kernel techniques we developed in 2016 and made publicly available through the model optimization toolkit in the TensorFlow Lite library. Model quantization delivered a 4x compression with respect to the trained floating point models and a 4x speedup at run-time, enabling our RNN-T to run faster than real time speech on a single core. After compression, the final model is 80MB.

Our new all-neural, on-device Gboard speech recognizer is initially being launched to all Pixel phones in American English only. Given the trends in the industry, with the convergence of specialized hardware and algorithmic improvements, we are hopeful that the techniques presented here can soon be adopted in more languages and across broader domains of application.

Acknowledgements:
Raziel Alvarez, Michiel Bacchiani, Tom Bagby, Françoise Beaufays, Deepti Bhatia, Shuo-yiin Chang, Zhifeng Chen, Chung-Chen Chiu, Yanzhang He, Alex Gruenstein, Anjuli Kannan, Bo Li, Wei Li, Qiao Liang, Ian McGraw, Patrick Nguyen, Ruoming Pang, Rohit Prabhavalkar, Golan Pundak, Kanishka Rao, David Rybach, Tara Sainath, Haşim Sak, June Yuan Shangguan, Matt Shannon, Mohammadinamul Sheik, Khe Chai Sim, Gabor Simko, Trevor Strohman, Mirkó Visontai, Ron Weiss, Yonghui Wu, Ding Zhao, Dan Zivkovic, and Yu Zhang.

Microsoft launches business school focused on AI strategy, culture and responsibility

In recent years, some of the world’s fastest growing companies have deployed artificial intelligence to solve specific business problems. In fact, according to new market research from Microsoft on how AI will change leadership, these high-growth companies are more than twice as likely to be actively implementing AI as lower-growth companies.

What’s more, high-growth companies are further along in their AI deployments, with about half planning to use more AI in the coming year to improve decision making compared to about a third of lower growth companies. Still, less than two in 10 of even high-growth companies are integrating AI across their operations, the research found.

“There is a gap between what people want to do and the reality of what is going on in their organizations today, and the reality of whether their organization is ready,” said Mitra Azizirad, corporate vice president for AI marketing at Microsoft in Redmond, Washington.

“Developing a strategy for AI extends beyond the business issues,” she explained. “It goes all the way to the leadership, behaviors and capabilities required to instill an AI-ready culture in your organization.”

On the road to developing a strategy, executives and other business leaders are often stalled by questions about how and where to begin implementing AI across their companies; the cultural changes that AI requires companies to make; and how to build and use AI in ways that are responsible, protect privacy and security, and comply with government rules and regulations.

Today, Azizirad and her team are launching Microsoft’s AI Business School to help business leaders navigate these questions. The free, online course is a master class series that aims to empower business leaders to lead with confidence in the age of AI.

 

YouTube Video

Focus on strategy, culture and responsibility

AI Business School course materials include brief written case studies and guides, plus videos of lectures, perspectives and talks that busy executives can access in small doses when they have time. A series of short introductory videos provide an overview of the AI technologies driving change across industries, but the bulk of the content focuses on managing the impact of AI on company strategy, culture and responsibility.

“This school is a deep dive into how you develop a strategy and identify blockers before they happen in the implementation of AI in your organization,” said Azizirad.

The business school complements other AI learning initiatives across Microsoft, including the developer-focused AI School and the Microsoft Professional Program for Artificial Intelligence, which provides job-ready skills and real-world experience to engineers and others looking to improve their skills in AI and data science.

Unlike these other initiatives, AI Business School is non-technical and designed to get executives ready to lead their organizations on a journey of AI transformation, according to Azizirad.

Nick McQuire, an analyst who covers artificial intelligence for CCS Insight, said more than 50 percent of the companies his firm has surveyed are already either researching, trialing or implementing specific projects with AI and machine learning, but very few are using AI across their organization and identifying business opportunities and problems that AI can address.

“That’s because there’s limited understanding in the business community about what AI is, what it can do and, ultimately, what are the applications,” he said. “Microsoft is trying to fill that gap.”

Mitra Azizirad stands in front of a Microsoft building looking into the camera Mitra Azizirad, corporate vice president for AI marketing. Photo by Microsoft.

Teaching by example

INSEAD, a graduate business school with campuses in Europe, Asia and the Middle East, partnered with Microsoft to build the AI Business School’s strategy module, which includes case studies about companies across many industries that have successfully transformed their businesses with AI.

For example, a case study on Jabil describes how one of the world’s largest manufacturing solutions providers was able to reduce overhead costs and increase production line quality by using AI to check electronic parts as they are manufactured, freeing up employees to focus on value added activities that machines are unable to do.

“There is still a lot of work that has got to have the human capital piece in it, especially if it is not something that lends itself to standardized processes,” explained Gary Cantrell, senior vice president and chief information officer for Jabil.

A key to implementing AI, Cantrell added, was the leadership team’s focus on clearly communicating to employees the company’s strategy around AI – to eliminate routine, repetitive activities in order to free them up to focus on activities that cannot be automated.

“If they are guessing or they are speculating, it is undoubtedly going to become counterproductive at some point,” he said. “So, the better job you do at keeping the team glued together with where you are going, the better the adoption will be and the faster it will be.”

Prepping an AI-ready culture

The culture and responsibility modules of AI Business School also place a core focus on data. After all, companies that successfully embrace AI need to openly share data across departments and business functions, explained Azizirad, and make sure all employees can participate in the development and implementation of data-driven AI applications.

“You need to start out with an open approach to how the data of an organization is going to be used, which is the foundation of AI, to get the results that you are banking on,” she said, adding that successful leaders foster an inclusive approach to AI that brings different roles together and breaks down data silos.

To illustrate the point, the Microsoft AI Business School surfaces a case study from Microsoft’s marketing team, which wanted to use AI to better score leads for the sales team to pursue. To build the solution, marketing employees partnered with data scientists to create machine learning models that weigh thousands of variables to score leads. The collaboration brought together marketing employees’ knowledge on lead quality with the machine learning expertise of data scientists.

“In the case of AI and in the case of culture, the people closest to the business problem you are trying to solve really need to be involved,” said Azizirad, adding that the sales team is embracing the lead-scoring model because they trust it will produce high-quality leads.

AI and responsibility

Building trust also comes from developing and deploying AI systems in a responsible manner, an area that Microsoft’s market research has found resonates with business leaders. Among high-growth companies, the research found, the more leaders know about AI, the more they recognize that they need to make sure the AI is deployed responsibly.

The AI Business School module on the implications of responsible AI showcases Microsoft’s own work in this area. Course materials include real-world examples in which leaders at Microsoft learned lessons such as the need to safeguard AI systems against malicious attacks and the need for systems to detect bias in datasets used to train models.

“Over time, as companies become operationally dependent on these machine learning algorithms and models that they built, there’s going to be much more focus on governance,” said McQuire, the CCS Insight analyst.

Related:

John Roach writes about Microsoft research and innovation. Follow him on Twitter.

The post Microsoft launches business school focused on AI strategy, culture and responsibility appeared first on The AI Blog.