HN user

bsprings

233 karma
Posts83
Comments10
View on HN
devblogs.nvidia.com 8y ago

Fast INT8 Inference for Autonomous Vehicles

bsprings
1pts0
devblogs.nvidia.com 8y ago

RESTful Inference with the TensorRT Container and Nvidia GPU Cloud

bsprings
2pts0
devblogs.nvidia.com 8y ago

How Jet.com Built a GPU-Accelerated Fulfillment Engine with F# and CUDA

bsprings
7pts0
twitter.com 8y ago

Malware Detection in Executables Using Neural Networks

bsprings
2pts1
devblogs.nvidia.com 8y ago

DeepStream: Next-Generation Video Analytics for Smart Cities

bsprings
3pts0
devblogs.nvidia.com 8y ago

Programming Tensor Cores in CUDA 9

bsprings
2pts0
devblogs.nvidia.com 8y ago

Mixed-Precision Deep Learning Training: Perf of FP16 with Accuracy of FP32

bsprings
1pts0
devblogs.nvidia.com 8y ago

Training AI for Self-Driving Cars: Challenges of Scale

bsprings
2pts0
devblogs.nvidia.com 8y ago

Cooperative Groups: Flexible CUDA Thread Programming

bsprings
2pts0
devblogs.nvidia.com 8y ago

Accelerated XGBoost Gradient Boosting on GPUs

bsprings
3pts0
devblogs.nvidia.com 8y ago

SigOpt: Deep Learning Hyperparameter Optimization with Competing Objectives

bsprings
6pts1
devblogs.nvidia.com 9y ago

2X Faster Embedded Small Batch Inference with Jetpack 3.1

bsprings
1pts0
devblogs.nvidia.com 9y ago

Deep Learning for Automated Driving with Matlab

bsprings
1pts0
www.reddit.com 9y ago

AI Co-Pilot: Driver Assistance via RNNs for Dynamic Facial Analysis

bsprings
9pts1
devblogs.nvidia.com 9y ago

GOAI: Open, GPU-Accelerated Data Analytics (demo Walkthrough)

bsprings
3pts0
devblogs.nvidia.com 9y ago

Explaining how end-to-end deep learning steers a self-driving car

bsprings
6pts0
devblogs.nvidia.com 9y ago

Tutorial: Manipulating Celebrity Faces with GANs

bsprings
2pts0
devblogs.nvidia.com 9y ago

Photo Editing with GANs: Intro to Generative Adversarial Networks

bsprings
2pts0
devblogs.nvidia.com 9y ago

Tutorial: Photo Editing with Generative Adversarial Networks (part 1)

bsprings
2pts0
devblogs.nvidia.com 9y ago

Caffe2: Get Started with new deep learning framework from Facebook

bsprings
2pts0
devblogs.nvidia.com 9y ago

Nvidia DGX-1 Deep Learning White Paper

bsprings
2pts0
devblogs.nvidia.com 9y ago

Personalized Aesthetics: Recording the Visual Mind Using Machine Learning

bsprings
2pts0
devblogs.nvidia.com 9y ago

Nvidia Jetson TX2: 2x perf & efficiency for IoT, robots, autonomous vehicles

bsprings
1pts0
devblogs.nvidia.com 9y ago

Nvidia DIGITS Assists Alzheimer's Disease Prediction

bsprings
6pts0
devblogs.nvidia.com 9y ago

Beyond GPU Memory Limits with Unified Memory on Pascal

bsprings
1pts0
devblogs.nvidia.com 9y ago

Image Segmentation using DIGITS 5

bsprings
3pts0
devblogs.nvidia.com 9y ago

New Compiler Features in CUDA 8 (faster; new C++ features)

bsprings
2pts0
devblogs.nvidia.com 9y ago

Deep Learning in Aerial Systems Using Jetson

bsprings
2pts0
devblogs.nvidia.com 9y ago

Mixed-Precision Programming with CUDA 8

bsprings
2pts0
devblogs.nvidia.com 9y ago

CUDA 8 Released: Pascal, Unified Memory, Mixed Precision

bsprings
3pts0

When I originally wrote the post in 2013, the GPU compilation part of Numba was a product (from Anaconda Inc., nee Continuum Analytics) called NumbaPro. It was part of a commercial package called Anaconda Accelerate that also included wrappers for CUDA libraries like cuBLAS, as well as MKL acceleration on the CPU.

Continuum gradually open sourced all of it (and changed their name to Anaconda). The compiler functionality is all open source within Numba. Most recently they released the CUDA library wrappers in a new open source package called pyculib.

Some other minor things changed, such as what you need to import. Also, the autojit and cudajit functionality is a bit better at type inference, so you don't have to annotate all the types to get it to compile.

We thought it was a good idea to update the post in light of all the changes.

(Post author here.) Yes, FP16 has been supported in NVIDIA GPUs as a texture format "forever" -- since before it was incorporated into the IEEE 754 standard. Indeed what is new in GP100 is hardware ALU support (and note denormals are full speed, which is even more important for lower precision formats).

FYI, nvprof works quite well with MPI, as described in this blog post by Jiri Kraus: http://devblogs.nvidia.com/parallelforall/cuda-pro-tip-profi...

To use nvprof with MPI, you just need to ensure nvprof is available on the cluster nodes and run it as your mpirun target, e.g. “mpirun ... nvprof ./my_mpi_program"

You can have it dump its output to files that the NVIDIA Visual Profiler (NVVP) is able to load. You can even load the output from multiple MPI ranks into NVVP to visualize them on the same timeline, making it easier to spot issues.

"unlike a real datacenter, it's only good for floats." That's actually not true. You'll find integer throughput is rather high on GPUs also. And the memory bandwidth is very high too. "miniature 4-function calculators": make that 5, where the fifth function is a really fast special function unit that can do very fast sin, cos, sqrt, 1/sqrt, and many other functions.

Hi varelse, can you tell me more about your profiling use case? nvprof should support MPI profiling scenarios, but perhaps yours is different. I'd love to know details so I can help improve the product. Feel free to contact me at first initial last name at nvidia.com (name is Mark Harris).