Skip to main content

Blog

Learn About Our Meetup

5000+ Members

MEETUPS

LEARN, CONNECT, SHARE

Join our meetup, learn, connect, share, and get to know your Toronto AI community. 

JOB POSTINGS

INDEED POSTINGS

Browse through the latest deep learning, ai, machine learning postings from Indeed for the GTA.

CONTACT

CONNECT WITH US

Are you looking to sponsor space, be a speaker, or volunteer, feel free to give us a shout.

Author: torontoai

[D] CIFAR-10 equivalents in video classification (action recognition)

So I’ve been working on some stuff with CNNs trained on CIFAR-10, and I’m interested in seeing if what I’ve been working on scales to video, i.e. with 3D CNNs.

Unfortunately, I haven’t been able to find any CIFAR-10 equivalents for videos. Most of the well-recognized datasets (UCF-101, Kinetics, Moments in Time, etc) are absolutely massive and correspond more with ImageNet.

Does anyone know of recognized and (relatively) small-scale video datasets that are good for sanity-checking and won’t require a massive overhead to work with? I have looked at UCF-11, but it is labeled by the authors as an “incredibly challenging” dataset, which is not what I’m looking for. Thanks for your help!

submitted by /u/ukhan_actual
[link] [comments]

[D] Questions about TFX after Google I/O’10 Talk: Machine Learning Pipelines and Model Understanding

TensorFlow Extended: Machine Learning Pipelines and Model Understanding (Google I/O’19)

1) Does anyone know of where someone might find an example repo showing functioning code of all TFX components? Maybe along the lines of a github repo or smth? E.g using example_gen, statistics_gen, etc.

2) In terms of deployment, is there anything to specifically look out for (esp with Tf2.0)? E.g eager mode doesn’t work like graph mode in production or keras estimators might be faster or slower than a low-level op estimator? Would really appreciate some guidance. Didn’t exactly find a central place

submitted by /u/iamquah
[link] [comments]

An End-to-End AutoML Solution for Tabular Data at KaggleDays

Machine learning (ML) for tabular data (e.g. spreadsheet data) is one of the most active research areas in both ML research and business applications. Solutions to tabular data problems, such as fraud detection and inventory prediction, are critical for many business sectors, including retail, supply chain, finance, manufacturing, marketing and others. Current ML-based solutions to these problems can be achieved by those with significant ML expertise, including manual feature engineering and hyper-parameter tuning, to create a good model. However, the lack of broad availability of these skills limits the efficiency of business improvements through ML.

Google’s AutoML efforts aim to make ML more scalable and accelerate both research and industry applications. Our initial efforts of neural architecture search have enabled breakthroughs in computer vision with NasNet, and evolutionary methods such as AmoebaNet and hardware-aware mobile vision architecture MNasNet further show the benefit of these learning-to-learn methods. Recently, we applied a learning-based approach to tabular data, creating a scalable end-to-end AutoML solution that meets three key criteria:

  • Full automation: Data and computation resources are the only inputs, while a servable TensorFlow model is the output. The whole process requires no human intervention.
  • Extensive coverage: The solution is applicable to the majority of arbitrary tasks in the tabular data domain.
  • High quality: Models generated by AutoML has comparable quality to models manually crafted by top ML experts.

To benchmark our solution, we entered our algorithm in the KaggleDays SF Hackathon, an 8.5 hour competition of 74 teams with up to 3 members per team, as part of the KaggleDays event. The first time that AutoML has competed against Kaggle participants, the competition involved predicting manufacturing defects given information about the material properties and testing results for batches of automotive parts. Despite competing against participants thats were at the Kaggle progression system Master level, including many who were at the GrandMaster level, our team (“Google AutoML”) led for most of the day and ended up finishing second place by a narrow margin, as seen in the final leaderboard.

Our team’s AutoML solution was a multistage TensorFlow pipeline. The first stage is responsible for automatic feature engineering, architecture search, and hyperparameter tuning through search. The promising models from the first stage are fed into the second stage, where cross validation and bootstrap aggregating are applied for better model selection. The best models from the second stage are then combined in the final model.

The workflow for the “Google AutoML” team was quite different from that of other Kaggle competitors. While they were busy with analyzing data and experimenting with various feature engineering ideas, our team spent most of time monitoring jobs and and waiting for them to finish. Our solution for second place on the final leaderboard required 1 hour on 2500 CPUs to finish end-to-end.

After the competition, Kaggle published a public kernel to investigate winning solutions and found that augmenting the top hand-designed models with AutoML models, such as ours, could be a useful way for ML experts to create even better performing systems. As can be seen in the plot below, AutoML has the potential to enhance the efforts of human developers and address a broad range of ML problems.

Potential model quality improvement on final leaderboard if AutoML models were merged with other Kagglers’ models. “Erkut & Mark, Google AutoML”, includes the top winner “Erkut & Mark” and the second place “Google AutoML” models. Erkut Aykutlug and Mark Peng used XGBoost with creative feature engineering whereas AutoML uses both neural network and gradient boosting tree (TFBT) with automatic feature engineering and hyperparameter tuning.

Google Cloud AutoML Tables
The solution we presented at the competitions is the main algorithm in Google Cloud AutoML Tables, which was recently launched (beta) at Google Cloud Next ‘19. The AutoML Tables implementation regularly performs well in benchmark tests against Kaggle competitions as shown in the plot below, demonstrating state-of-the-art performance across the industry.

Third party benchmark of AutoML Tables on multiple Kaggle competitions

We are excited about the potential application of AutoML methods across a wide range of real business problems. Customers have already been leveraging their tabular enterprise data to tackle mission-critical tasks like supply chain management and lead conversion optimization using AutoML Tables, and we are excited to be providing our state-of-the-art models to solve tabular data problems.

Acknowledgements
This project was only possible thanks to Google Brain team members Ming Chen, Da Huang, Yifeng Lu, Quoc V. Le and Vishy Tirumalashetty. We also thank Dawei Jia, Chenyu Zhao and Tin-yun Ho from the Cloud AutoML Tables team for great infrastructure and product landing collaboration. Thanks to Walter Reade, Julia Elliott and Kaggle for organizing such an engaging competition.

[P] Language model (GPT-2) fine-tuned on ‘The Office’ dialogues

I fine-tuned the newly released OpenAI GPT-2 345M model on all dialogues (Season 1 to 9, around 3MB) of ‘The Office’ and the model starts to acquire some humour, at least when reading it through the lens of the specific characters.
Some of my favourites:

“Michael: What is the most important thing for a company? Is it the cash flow? Is it the inventory? Oh, it’s the closets? The closets are the best.”
———————————————————————————————————————————

“Michael: All right, time for another round of questions. First, how old are you?”
“Phyllis: I’m 58. “
“Oscar: 58? “
“Kevin: Hey, hey, that’s old, right? Can I get a raise?”

———————————————————————————————————————————

“Andy: This book’s worth more to me than any number of my books. I mean, I don’t care. One of them is a good book, the other is a bad book, and the next one is a joke. All right? So… my money’s on one of the two, or zero, as they say in French, so that’s fine. “

I published some more here and will add some after training the model for a bit longer:

–> https://twitter.com/fateful_by

Some things I already noticed:

  • It got really dark at some points, e.g. “Darryl: I’m going to take a little piece of ice. I’m going to stab it into your skull.” or “Dwight: [in Andy’s car] You know, I’ve always wanted to be a detective. So, when I come back, I’m gonna sniff out leads. And if I catch a murderer, I’m gonna go blow my brains out. Now, I don’t know how that works. But I think I might be the best person to go out there and catch one of those serial killers. “
  • Temperature setting above 0.7 gives some really crazy outputs, e.g. “Andy: Stanley Hudson! [throws tie to Dwight, who throws it back] Dwight: Sorry, I didn’t mean to upset you. Ralph knows best. [to Creed] What’s gay is who you know and what’s deb (back to Oscar) You can’t work here. You’re a too busy hand. [back to Toby] You were fired for stealing the pen, right? Toby: Yesss! Bring it on!

submitted by /u/CYHSM
[link] [comments]

[P] Finding structures vulnerable to disaster in street view imagery

[P] Finding structures vulnerable to disaster in street view imagery

My company is supporting the World Bank’s program focused on resilient housing. Housing “resiliency” is a problem in poor urban areas around the world as these homes are often not up to code (e.g., because families are doing their own house construction or good materials aren’t available). This makes the homes (and the families inside) vulnerable to disasters like earthquakes and hurricanes. Many governments in these countries have resources to retrofit the houses for safety, but it’s inefficient for structural engineers to walk up and down neighborhoods to identify the vulnerable homes that need fixing.

We just finished some pilot work to identify building features that could confer risk. The features we detected here were

  1. building material
  2. whether or not a building looked designed by an architect
  3. and whether or not construction appeared complete.

This is definitely more application focused, so we just used TF’s object detection API under the hood w/ an SSD backbone and trained on ~7000 labeled street view images. The fun part was relating street view detections to an overhead building footprint map — something I talk about a lot in the blog. (Does anyone know how Google or other tech companies do this?) We’re working on releasing that training data if anyone else is interested in this problem space.

One related question: does anyone know if TF OD API will change in TF 2.0?

Edit: adding gif example

Processing gif afv9d79i78x21…

submitted by /u/wronk17
[link] [comments]

[P] SOD – An Embedded Computer Vision and Machine Learning Library

Project Homepage

sod.pixlab.io

Purpose

SOD is an embedded, modern cross-platform computer vision and machine learning software library that expose a set of APIs for deep-learning, advanced media analysis & processing including real-time, multi-class object detection and model training on embedded systems with limited computational resource and IoT devices.

Docs

edit: Markdown

submitted by /u/histoire_guy
[link] [comments]