Skip to main content

Blog

Learn About Our Meetup

5000+ Members

MEETUPS

LEARN, CONNECT, SHARE

Join our meetup, learn, connect, share, and get to know your Toronto AI community. 

JOB POSTINGS

INDEED POSTINGS

Browse through the latest deep learning, ai, machine learning postings from Indeed for the GTA.

CONTACT

CONNECT WITH US

Are you looking to sponsor space, be a speaker, or volunteer, feel free to give us a shout.

Author: torontoai

By the Book: AI Making Millions of Ancient Japanese Texts More Accessible

Natural disasters aren’t just threats to people and buildings, they can also erase history — by destroying rare archival documents. As a safeguard, scholars in Japan are digitizing the country’s centuries-old paper records, typically by taking a scan or photo of each page.

But while this method preserves the content in digital form, it doesn’t mean researchers will be able to read it. Millions of physical books and documents were written in an obsolete script called Kuzushiji, legible to fewer than 10 percent of Japanese humanities professors.

“We end up with billions of images which will take researchers hundreds of years to look through,” said Tarin Clanuwat, researcher at Japan’s ROIS-DS Center for Open Data in the Humanities. “There is no easy way to access the information contained inside those images yet.”

Extracting the words on each page into machine-readable, searchable form takes an extra step: transcription, which can be done either by hand or through a computer vision method called optical character recognition, or OCR.

Clanuwat and her colleagues are developing a deep learning OCR system to transcribe Kuzushiji writing — used for most Japanese texts from the 8th century to the start of the 20th — into modern Kanji characters.

Clanuwat said GPUs are essential for both training and inference of the AI.

“Doing it without GPUs would have been inconceivable,” she said. “GPU not only helps speed up the work, but it makes this research possible.”

Parsing a Forgotten Script

Before the standardization of the Japanese language in 1900 and the advent of modern printing, Kuzushiji was widely used for books and other documents. Though millions of historical texts were written in the cursive script, just a few experts can read it today.

Only a tiny fraction of Kuzushiji texts have been converted to modern scripts — and it’s time-consuming and expensive for an expert to transcribe books by hand. With an AI-powered OCR system, Clanuwat hopes a larger body of work can be made readable and searchable by scholars.

She collaborated on the OCR system with Asanobu Kitamoto from her research organization and Japan’s National Institute of Informatics, and Alex Lamb of the Montreal Institute for Learning Algorithms. Their paper was accepted in 2018 to the Machine Learning for Creativity and Design workshop at the prestigious NeurIPS conference.

Using a labeled dataset of 17th to 19th century books from the National Institute of Japanese Literature, the researchers trained their deep learning model on NVIDIA GPUs, including the TITAN Xp. Training the model took about a week, Clanuwat said, but “would be impossible” to train on CPU.

Kuzushiji has thousands of characters, with many occurring so rarely in datasets that it is difficult for deep learning models to recognize them. Still, the average accuracy of the researchers’ KuroNet document recognition model is 85 percent — outperforming prior models.

The newest version of the neural network can recognize more than 2,000 characters. For easier documents with fewer than 300 character types, accuracy jumps to about 95 percent, Clanuwat said. “One of the hardest documents in our dataset is a dictionary, because it contains many rare and unusual words.”

One challenge the researchers faced was finding training data representative of the long history of Kuzushiji. The script changed over the hundreds of years it was used, while the training data came from the more recent Edo period.

Clanuwat hopes the deep learning model could expand access to Japanese classical literature, historical documents and climatology records to a wider audience.

The post By the Book: AI Making Millions of Ancient Japanese Texts More Accessible appeared first on The Official NVIDIA Blog.

Model-Based Reinforcement Learning from Pixels with Structured Latent Variable Models

Imagine a robot trying to learn how to stack blocks and push objects using
visual inputs from a camera feed. In order to minimize cost and safety
concerns, we want our robot to learn these skills with minimal interaction
time, but efficient learning from complex sensory inputs such as images is
difficult. This work introduces SOLAR, a
new model-based reinforcement learning (RL) method that can learn skills –
including manipulation tasks on a real Sawyer robot arm – directly from
visual inputs with under an hour of interaction. To our knowledge, SOLAR is the
most efficient RL method for solving real world image-based robotics tasks.




Our robot learns to stack a Lego block and push a mug onto a coaster with only
inputs from a camera pointed at the robot. Each task takes an hour or less of
interaction to learn.

Continue reading

[D] Successfully Deploying GPT-2 in a Flask Web App

Has anyone been able to successfully, smoothly deploy GPT-2 in a Flask web app? If you have, would you mind pointing me in the direction of resources to help accomplish this myself?

I’ve seen a few instances of GPT-2 deployed with UI allowing users to generate text on their own, but the performance has been much shoddier and more erratic than what I’m capable of producing on my own system (sampling is either straight up bugged, or spits out fairly repetitive mush).

I’m in the later stages of building a platform that presents educational materials relating to language models, and some of machine learning’s role in NLP tasks, from understanding to generation, hitting upon super entry-level concepts, recent architectural innovation, and some social/societal ramifications of the state of the art, etc. All in all, it’s aimed at relative beginners and curious enthusiasts, but also includes plenty of resources for more intermediate and advanced users.

I’d really like to incorporate an interactive space for my users to experiment with GPT-2 firsthand. Website and everything is already built out, in Flask, just sort of struggling with some devops and deploy stuff.

submitted by /u/Extension_Juggernaut
[link] [comments]

[D] Machine Learning – WAYR (What Are You Reading) – Week 63

This is a place to share machine learning research papers, journals, and articles that you’re reading this week. If it relates to what you’re researching, by all means elaborate and give us your insight, otherwise it could just be an interesting paper you’ve read.

Please try to provide some insight from your understanding and please don’t post things which are present in wiki.

Preferably you should link the arxiv page (not the PDF, you can easily access the PDF from the summary page but not the other way around) or any other pertinent links.

Previous weeks :

1-10 11-20 21-30 31-40 41-50 51-60 61-70
Week 1 Week 11 Week 21 Week 31 Week 41 Week 51 Week 61
Week 2 Week 12 Week 22 Week 32 Week 42 Week 52 Week 62
Week 3 Week 13 Week 23 Week 33 Week 43 Week 53
Week 4 Week 14 Week 24 Week 34 Week 44 Week 54
Week 5 Week 15 Week 25 Week 35 Week 45 Week 55
Week 6 Week 16 Week 26 Week 36 Week 46 Week 56
Week 7 Week 17 Week 27 Week 37 Week 47 Week 57
Week 8 Week 18 Week 28 Week 38 Week 48 Week 58
Week 9 Week 19 Week 29 Week 39 Week 49 Week 59
Week 10 Week 20 Week 30 Week 40 Week 50 Week 60

Most upvoted papers two weeks ago:

/u/gatapia: Unsupervised learning by competing hidden units

/u/Consistent_Size: “A tutorial on subspace clustering (pdf)”

/u/vlanins: Poverty Mapping Using Convolutional Neural Networks Trained on High and Medium Resolution Satellite Images, With an Application in Mexico

Besides that, there are no rules, have fun.

submitted by /u/ML_WAYR_bot
[link] [comments]

[P] New in DeOldify: Smooth Colorization of Video! Courtesy of one weird trick- NoGAN

[P] New in DeOldify: Smooth Colorization of Video! Courtesy of one weird trick- NoGAN

Hello again! I just realized after a few weeks of sitting on this that I didn’t tell you guys about something really cool I’ve been working on! Maybe I should do that.

A few months ago I posted about DeOldify- my pet project for colorizing and restoring old photos. Well now that includes videos as well! Here’s a demo I showed at Facebook’s F8 conference:

https://reddit.com/link/bq8gji/video/9tu80qji41z21/player

And here’s the talk at F8: https://www.facebook.com/FacebookforDevelopers/videos/340167420019712/

And here’s the article I wrote with Jeremy Howard of FastAI and Uri Manor of Salk Institute: https://www.fast.ai/2019/05/03/decrappify/

Anyway, the gist is that we (FastAI and I) developed this one weird trick called NoGAN to achieve this. That basically consists of this (slide from F8):

https://i.redd.it/midd5z9v41z21.jpg

The pretraining is using basic perceptual loss (or “feature loss”) for the generator. This gets you the benefits of GANs, without the problems, basically. Hence, smooth video!

The progression of training looks like this (sweet spot is at 1.4%, then it goes too far from there and gets weird with the orange skin):

https://reddit.com/link/bq8gji/video/34o1egby41z21/player

Anyway, that’s the gist. You can read more in the links above (readme for github project also has good details on NoGAN). Oh by the way, NoGAN also works on super resolution, and I suspect for most image to image tasks as well as perhaps even non-image tasks.

https://i.redd.it/z0tlr0zz41z21.png

submitted by /u/MyMomSaysImHot
[link] [comments]