Skip to main content

Blog

Learn About Our Meetup

5000+ Members

MEETUPS

LEARN, CONNECT, SHARE

Join our meetup, learn, connect, share, and get to know your Toronto AI community. 

JOB POSTINGS

INDEED POSTINGS

Browse through the latest deep learning, ai, machine learning postings from Indeed for the GTA.

CONTACT

CONNECT WITH US

Are you looking to sponsor space, be a speaker, or volunteer, feel free to give us a shout.

Author: torontoai

[P] I’m working on a program to help teach Q-Learning

During my last Machine Learning course I was working with Q-Learning and made a small program that allows users to create maps for a Q-Learning agent to solve. I am using a game engine called Godot for the visuals and coded in Godot’s custom language GDScript. My teacher felt that I should post this online for others to use to help learn Q-Learning. I plan to add menu options to provide a detailed walkthrough of how the program works and possibly other Q-Learning example projects.

Github: https://github.com/bioa10/QTable

submitted by /u/bioa10
[link] [comments]

[D] Should I bring up ethics in my thesis?

For my master’s thesis, I’m working on a model that could potentially be misused with malicious intents. It’s getting a common theme lately.

I’ve got to admit that I don’t have a strong sense of morals and that I work on that project mostly because I love to see machine learning in action, and I don’t really care of the actual applications of my model. I intend to make my trained models public and the code open source. While I doubt any of it will actually cause havoc, I feel like the topic of responsible AI is looming to make a paragraph in my thesis. I don’t really like it, because I’d prefer to entirely avoid the subject. So far, I’ve only written

(…) is attractive for a range of applications be they useful, merely a matter of customization, or mischievous.

I feel like it would make me look bad to entirely avoid the topic when I’m planning to make my work publicly available.

submitted by /u/Valiox
[link] [comments]

[N] ICML 2019 Accepted Paper Stats

[N] ICML 2019 Accepted Paper Stats

I recently compiled some stats and figures regarding the accepted papers at this years International Conference on Machine Learning (ICML). All data is taken from https://icml.cc/Conferences/2019/AcceptedPapersInitial

Here you go:

Top contributing institutes @ ICML 2019 according to the number of papers with at least one author affiliated with this institute. Ordered according to the total number of papers (followed by number of first and last author papers).

Here are some figures about top contributing authors.

Top contributing authors @ ICML 2019 according to the total number of papers

Top contributing authors @ ICML 2019 according to their number of first author papers.

Top contributing authors @ ICML 2019 according to their number of last author papers.

Finally, some stats for top contributing institutes sorted by their relative contribution (i.e. how many authors on a paper a actually from this institute).

https://i.redd.it/o1783e2oeix21.png

I compiled these stats for my employer, the Robert Bosch GmbH. If you’re interested in Bosch and what we’re doing in AI and ML, check out https://www.bosch-ai.com/

Disclaimer

Cleaning the website data, in particular the affiliations is a tedious, manual process, since many different and not necessarily unambiguous notations and abbreviations exist for many institutes. I tried my best to merge affiliation names into distinct institute buckets. However, there might be errors in the data, leading to single papers not being counted for a specific institute. Same applies for author names, which are NOT manually merged due to the large number of authors. Only identical author names have been associated across different publications.

submitted by /u/AndreasDoerr
[link] [comments]

[N] Using Prodmodel to speed up data science development and productionization

https://github.com/prodmodel/prodmodel

I built a tool which keeps track of all code and data deps of your data science project. It caches partial results and can figure out if a particular output (model, transformed data, code library) has to be recomputed before the fact. This can save huge amounts of time during an iterative development process.

Setting up usage is similar to build systems like Bazel or Make (https://github.com/prodmodel/prodmodel/blob/master/example/build.py).

Feedback, users, contributors and constructive criticism is welcome!

submitted by /u/gsvigruha
[link] [comments]

[P] Not sure if this is too silly for this reddit but I used GPT-2 345M to recreate the famous (but fake) “I forced a bot to write an Olive Garden commercial” tweet, but I fine-tuned GPT-2 on Star Trek scripts first…

Link To The Post:

https://iforcedabot.com/i-forced-a-bot-to-watch-over-1000-of-star-trek-episodes-and-then-asked-it-to-write-1000-olive-garden-commercials/

What is this?

About a year ago, a tweet went viral purporting to be written by a bot: I forced a bot to watch over 1,000 hours of Olive Garden commercials and then asked it to write an Olive Garden commercial of its own. Here is the first page.

It was a funny script but it was written by a comedy writer, not a bot. This was obvious to some people, but not others, and enough people that it was real that sites wrote articles debunking it.

I’ve had this in the back of mind for while. It used to be obvious that scripts like this were fake — anyone who had used an RNN network knew they could only hold a coherent thought for about two sentences, there’s no way it would produce something like that. But since that tweet GPT-2 came out and changed everything. I was looking for something to test fine-tuning the new GPT-345M on and picked the dumbest silliest option.

I can’t tell you exactly why I also also mixed this with Star Trek The Next Generation and Deep Space Nine scripts except that it I found it hilarious.

How it works

The training set is all Star Trek, but GPT-2 is shockingly good at writing lines appropriate for a waitress or restaurant scene that are 100% are not in the Star Trek script training set. I love how creative it can be since it has such a wide range of pre-baked knowledge.

Are these hand selected samples?

Nope. The title says 1,000 commercials but it’s actually over 30,000 commercials now — I couldn’t resist trying different prompts and tweaks to the training data to get better results. I didn’t filter the samples at all so it’s a wide mix of iterations, temperatures, learning rates, tweaked training sets, I pretty much just tossed everything up there. I was overall impressed with how well GPT-2 worked with both very little fine-tuning, and how it avoided overfitting with a lot more training. I trained these same Trek scripts with the smaller GPT-2 and had over-fitting problems eventually but I never hit that point with 345M. I suppose it’s possible it just trains a lot slower…

submitted by /u/JonathanFly
[link] [comments]

The AWS DeepRacer League virtual circuit is underway—win a trip to re:Invent 2019!

The competition is heating up in the AWS DeepRacer League, the world’s first global autonomous racing league, open to anyone. The first round is almost halfway home, now that 9 of the 21 stops on the summit circuit schedule are complete. Developers continue to build new machine learning skills and post winning times to the leaderboards. Here’s a quick round-up of the news from all of this week’s action.

The AWS DeepRacer virtual circuit launched on April 29. Developers of all skill levels can enter the league from anywhere in the world via the AWS DeepRacer console.

The first of six monthly tracks is the London Loop, and racing is well underway. As of May 8, 2019, the are 346 participants on the leaderboard, competing to be crowned the first champion of the virtual circuit and advance on an all-expenses-paid trip to re:Invent. Our current leader is Holly, with a time of 12.48 seconds. Twenty-three days remain, so there’s still time to get rolling into the online competition. There are prizes for the Top 10, and plenty of chances to win!

Current leaderboard standings:

Time remaining on the London Loop race:

On the Summit Circuit this week, the AWS DeepRacer League made stops in Madrid and London and crowned two new champions. They both advance on an all-expenses-paid trip to re:Invent 2019 in Las Vegas, Nevada.

First up was Madrid, the third city in Europe to host the AWS DeepRacer League. The crowd was energetic and the competitors eager to win. The top 3 took to the tracks 14 times between them.

Pedro, Javier, and David arrived at the AWS Summit together, with 27 models that they had been training together in the AWS DeepRacer 3D racing simulator. They had seen some good results in the virtual world. However, the first couple of runs on the track didn’t seem to deliver in the same way, with our champion Pedro posting an opening time of 40 seconds. They pulled together as a team, tuning and trying the different models they had built at home, and eventually began to see much better results.

In the following video, David shares their thoughts on strategy during the day.

With about two hours of racing left, and on his fourth attempt, Pedro was the lucky team member who took the top spot with a winning time of 9.36 seconds. His colleagues were not far behind, claiming the second and third spot. Pedro advances to the finals and is excited to work with his teammates on a strategy to take home the AWS DeepRacer League Championship Cup. Don’t worry, they both join him to take on the rest of the field!

And on to London, the hometown of the reigning AWS DeepRacer Champion, Rick Fish. Developers came to the expo hall at the AWS Summit, for a full day of racing on two tracks and the chance to win their trip to re:Invent 2019.

The day started strong with our eventual third-place finisher “breadcentric,” with a 13-second lap. New to machine learning, he brought his model to the AWS Summit and was ready to race as soon as the tracks opened at 8AM. The competition came in strong as competitors quickly started logging lap times under 10 seconds, including our eventual champion, Matt Camp. Matt works at Jigsaw XYZ, whose cofounder happens to be Rick Fish! Rick’s team at Jigsaw XYZ had been preparing for the London race since re:Invent and knew that the pressure would be on to win.

Matt had been working on his model at home and was eager to see how well it could perform. Matt’s friend and colleague Tony joined him. With only 1 hour to go, they were in second and third position on the podium, behind Raul, who had spent most of the day on top with a 9.01-second lap. The Jigsaw XYZ team took to the tracks one more time. In his final 2 minutes of racing, Matt clinched the title with an 8.9-second lap. Matt had no experience with machine learning before re:Invent 2018. He now heads back in 2019 to take on Rick Fish and rest of the field to win the AWS DeepRacer League Championship Cup.

The competition and excitement are certainly building in the AWS DeepRacer League. Developers of all skill levels get hands-on, learn, and put their machine learning skills to the ultimate test. Get started in the AWS DeepRacer League, either virtually or at the next summit near you. We have all the tools to get you started even if you have no machine learning experience, as well as resources to help you take on the challenge and win!

Coming soon, we share our best tips from the AWS DeepRacer team, so stay tuned.


About the Author

Alexandra Bush is a Senior Product Marketing Manager for AWS AI. She is passionate about how technology impacts the world around us and enjoys being able to help make it accessible to all. Out of the office she loves to run, travel and stay active in the outdoors with family and friends.

 

 

 

[D] How does your machine learning algorithm are indrustialized ?

Hello r/MachineLearning,

Disclaimer, I am a data engineer. I hope this post have its place here, I apologies in advanced if not.

What is your experience regarding the industrialization of your ML code ? Who do you collaborate with?

I will start : I’ve worked with data scientists to release in production machine learning pipeline. We have been improving our collaboration over the past year, and thus delivering more efficiently. First by drafting a blue print of a machine learning pipeline, then converging on common tools, finally by sharing our knowledges on our different skills.

Our shared tech stack is mainly: BigQuery, Apache Beam (to distribute the preprocessing), docker, tensorflow/Keras, Apache airflow.

This has lead to data scientists being more autonomous regarding scalability.

I think collaboration is the key.

EDIT : cleaning, removing medium link.

submitted by /u/schrute_dataeng
[link] [comments]