Skip to main content

Blog

Learn About Our Meetup

5000+ Members

MEETUPS

LEARN, CONNECT, SHARE

Join our meetup, learn, connect, share, and get to know your Toronto AI community. 

JOB POSTINGS

INDEED POSTINGS

Browse through the latest deep learning, ai, machine learning postings from Indeed for the GTA.

CONTACT

CONNECT WITH US

Are you looking to sponsor space, be a speaker, or volunteer, feel free to give us a shout.

Author: torontoai

[P] Tool to build GPT-2 textgen APIs scalable and free using Google Cloud Run

https://github.com/minimaxir/gpt-2-cloud-run

There have been a few posts here w/ interactive GPT-2 textgen models. I’ve built an open-source tool to help build APIs with GPT-2 (specifically, fine-tuned models on a new dataset) via gpt-2-simple and deploy them to Cloud Run, where the pricing works out to be effectively free unless you have huge spikes or constant requests.

I have also included a mini-Cloud Build tutorial to limit model downloading/uploading.

submitted by /u/minimaxir
[link] [comments]

[R] Function which takes as input two vectors of the same dimension from different spaces and produces a single latent vector

Hey!

I was working on an extreme multi-label classification task of which one of the modules is to create a function that would take a single document vector and label vector (both of 300 dimensions but created separately and so from different vector spaces) and would produce a new latent vector which would be binarily classifiable. My current naive approach is to simply concatenate the two vectors.

Could you please suggest some techniques, features or a general direction or topic to study. Any idea would be greatly appreciated as I have been at a total loss on how to proceed further.

Thank you!

submitted by /u/atif_hassan
[link] [comments]

[P] A Tutorial on Multi-Label Classification using Deep Learning

This blog post provides an elaborate introductory tutorial on creating Deep Learning models for Multi-Label Classification. The concept is explored by creating a neural network in Keras (using TensorFlow) that can assign multiple labels to different food items. The code for this project is available in GitHub and can also be accessed through Google Colab. You can checkout the project and the article here:

Article: https://blog.nanonets.com/multi-label-classification-using-deep-learning/

GitHub: https://github.com/thatbrguy/Multilabel-Classification

I would love to hear your thoughts and feedback about the same. Thanks!

submitted by /u/thatbrguy_
[link] [comments]

[Research] A collection of 1 Million Computer-Aided Design (CAD) Models for Geometric Deep Learning Research

https://medium.com/ai%C2%B3-theory-practice-business/abc-free-datasets-for-geometric-deep-learning-5e2995768b37

Abstract: We introduce ABC-Dataset, a collection of one million Computer-Aided Design (CAD) models for research of geometric deep learning methods and applications. Each model is a collection of explicitly parametrized curves and surfaces, providing ground truth for differential quantities, patch segmentation, geometric feature detection, and shape reconstruction.

submitted by /u/cdossman
[link] [comments]

AI of the Storm: Deep Learning Analyzes Atmospheric Events on Saturn

Saturn is many times the size of Earth, so it’s only natural that its storms are more massive — lasting months, covering thousands of miles, and producing lightning bolts thousands of times more powerful.

While scientists have access to galaxies of data on these storms, its sheer volume leaves traditional methods inadequate for studying the planet’s weather systems in their entirety.

Now AI is being used to launch into that trove of information. Researchers from University College London and the University of Arizona are working with data collected by NASA’s Cassini spacecraft, which spent 13 years studying Saturn before disintegrating in the planet’s atmosphere in 2017.

A recently published Nature Astronomy paper describes how the scientists’ deep learning model can reveal previously undetected atmospheric features on Saturn, and provide a clearer view of the planet’s storm systems at a global level.

In addition to providing new insights about Saturn, the AI can shed light on the behavior of planets both within and beyond our solar system.

“We have these missions that go around planets for many years now, and much of this data basically sits in an archive and it’s not being looked at,” said Ingo Waldmann, deputy director of the University College London’s Centre for Space Exoplanet Data. “It’s been difficult so far to look at the bigger picture of this global dataset because people have been analyzing data by hand.”

The researchers used an NVIDIA V100 GPU, the most advanced data center GPU, for both training and inference of their neural networks.

Parting the Clouds

Scientists studying the atmosphere of other planets take one of two strategies, Waldmann says. Either they conduct a detailed manual analysis of a small region of interest, which could take a doctoral student years — or they simplify the data, resulting in rough, low-resolution findings.

“The physics is quite complicated, so the data analysis has been either quite old-fashioned or simplistic,” Waldmann said. “There’s a lot of science one can do by using big data approaches on old problems.”

Thanks to the Cassini satellite, researchers have terabytes of data available to them. Primarily using unsupervised learning, Waldmann and Caitlin Griffith, his co-author from the University of Arizona, trained their deep learning model on data from the satellite’s mapping spectrometer.

This data is commonly collected on planetary missions, Waldmann said, making it easy to apply their AI model to study other planets.

The researchers saw speedups of 30x when training their deep learning models on a single V100 GPU compared to CPU. They’re now transitioning to using clusters of multiple GPUs. For inference, Waldmann said the GPU was around twice as fast as using a CPU.

Using the AI model, the researchers were able to analyze a months-long electrical storm that churned through Saturn’s southern hemisphere in 2008. Scientists had previously detected a bright ammonia cloud from satellite images of the storm — a feature more commonly spotted on Jupiter, but rarely seen on Saturn.

Waldmann and Griffith’s AI-analyzed data from this months-long electric storm on Saturn. Left image shows the planet in colors similar to how the human eye would see it, while the image on the right is color enhanced, making the storm stand out more clearly. (Image credit: NASA/JPL/Space Science Institute)

Waldmann and Griffith’s neural network found that the ammonia cloud visible by eye was just the tip of a “massive upwelling” of ammonia hidden under a thin layer of other clouds and gases.

“What you can see by eye is just the strongest bit of that ammonia feature,” Waldmann said. “It’s just the tip of the iceberg, literally. The rest is not visible by eye — but it’s definitely there.”

To Infinity and Beyond

For researchers like Waldmann, these findings are just the first step. Deep learning can provide planetary scientists for the first time with depth and breadth at once, producing detailed analyses that also cover vast geographic regions.

“It will tell you very quickly what the global picture is and how it all connects together,” said Waldmann. “Then researchers can go and look at individual spots that are interesting within a particular system, rather than blindly searching.”

A better understanding of Saturn’s atmosphere can help scientists analyze how our solar system behaves, and provide insights that can be extrapolated to planets around other stars.

Already, the researchers are extending their model to study features on Mars, Venus and Earth using transfer learning — which they were surprised to learn “works really well between planets.”

While Venus and Earth are almost identical in size, Venus has no global plate tectonics. In collaboration with the Observatoire de Paris, the team is starting a project to analyze Venus’s cloud structure and planetary surface to understand why the planet lacks tectonic plates.

Rather than atmospheric features, the researchers’ Mars project focuses on studying the planet’s surface. Data from the Mars Reconaissance Orbiter can create a global analysis that scientists can use to deduce where ancient water was most likely present, and to determine where the next Mars rover should land.

The underlying pattern recognition algorithm can be extended even further, Waldmann said. On Earth, it can be repurposed to spot rogue fishing vessels to preserve protected environments. And across the solar system on Jupiter, a transfer learning approach can train an AI model to analyze how the planet’s storms change over time.

Waldmann says there’s relatively easy access to training data — creating an open field of opportunities for researchers.

“This is the beautiful thing about planetary science,” he said. “All of the data for all of the planets is publicly available.”

Main image, captured in 2011, shows the largest storm observed on Saturn by the Cassini spacecraft. (Image credit NASA/JPL-Caltech/SSI)

The post AI of the Storm: Deep Learning Analyzes Atmospheric Events on Saturn appeared first on The Official NVIDIA Blog.

[P] Replication and Comparisons of Disentangled VAE

Hi r/MachineLearning,

With some friends, we replicated some of the most important disentangled VAE (beta-TCVAE, FAactorVAE, both versions of beta-VAE). The motivations were the following:

  • Have a modular framework easily extendable for any type of VAE
  • Separate the loss and the architecture to have a better understanding of what causes the improvements in new papers (many papers change both and it’s thus hard to disentangle – no puns intended – where the improvements come from).
  • Have a clean and easily extendable visualisation pipeline for VAEs
  • Replicate quantitative metrics to compare disentanglement

I hope you find it useful, and please let me know if there is anything you think would be useful to add (new disentangle VAE loss, add GAN for reconstruction loss, …)

https://github.com/YannDubs/disentangling-vae

PS: it’s my first post on reddit, please let me know if I’m doing something wrong.

submitted by /u/yannDubs
[link] [comments]

Citibank Data Science Interview Questions

Citicorp and Travelers’ Group has total assets of $1.7 trillion.

In America, Citibank is one of the four main firms that accounts for half of the nation’s total mortgages, and two-thirds of the total credit cards. Although this institution isn’t necessarily the largest in America, it is often considered to be the largest banking facility across the globe. Citibank serves a mass number of over 200 million customers across a span of 160 countries. Its main location functions through Citibank Europe, stationed in the Czech Republic. The Spend Tracker which is quite common at banks all over the world today was started by Citibank first. Citibank which provides such heterogeneous financial products showcases a wide variety of information on its spend tracker for its customers from sign-on bonuses, bonus amounts and expiration dates, bonus miles, and even how much you have to spend in order to get certain rewards. Such varied information across multiple products and across 200 million customers makes it one of the best companies for Data Scientists to work at.

Photo by Anthony Ginsbrook on Unsplash

Interview Process

The interview process starts with a phone interview. The phone interview is a basic Data Science Q&A interview. The phone interview is followed by an onsite interview. The onsite interview consists of interview with team leads, team members and SVPs. There may or may not be an online SQL assessment before the onsite. The SQL assessment is usually a difficult one.

Important Reading

Source: ML and Cognitive Computing

Data Science Related Interview Questions

  • Given a list of integers, find all combinations that sum to a given integer.
  • Segment a long string into a set of valid words using a dictionary. Return false if the string cannot be segmented. What is the complexity of your solution?
  • Write a SQL query to find the repeated items in a column.
  • How do you use the Q data structure to maintain the state in Spark.
  • Design a Trading system with high throughput and low latency.
  • How do you describe a financial planning process?
  • What would you prefer, being attacked by a giant chicken or 100 small ones?
  • What problems did you encounter in your project(resume based) and what are the solutions you did?
  • Explain your thesis in layman’s terms.
  • Find the second maximum value of a column in a Database table.

Reflecting on the Questions

The data science team at Citigroup uses Hadoop and Spark. They have a geographical diverse team located in the US, Europe and India. Their questions are a mix of questions related to coding, SQL, Systems Design, Hadoop and Spark. They are based on foundational and deep aspects of Data Science. If you work hard on your basics, you can surely land a job at one of the largest banks of the world!

Subscribe to our Acing AI newsletter, I promise not to spam and its FREE!

Acing AI Newsletter – Revue

Thanks for reading! 😊 If you enjoyed it, test how many times can you hit 👏 in 5 seconds. It’s great cardio for your fingers AND will help other people see the story.

The sole motivation of this blog article is to learn about Citibank and its technologies helping people to get into it. All data is sourced from online public sources. I aim to make this a living document, so any updates and suggested changes can always be included. Please provide relevant feedback.


Citibank Data Science Interview Questions was originally published in Acing AI on Medium, where people are continuing the conversation by highlighting and responding to this story.