Mesh-TensorFlow: Model Parallelism for Supercomputers (TF Dev Summit 19)

About Share Download Add to

Batch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming. However, batch-splitting suffers from problems including the inability to train very large models (due to memory constraints), high latency, and inefficiency at small batch sizes. All of these can be solved by more general distribution strategies (model-parallelism). Unfortunately, efficient mode

Share with your friends

Link:

Embed:

<iframe width="640" height="360" src="//myvideo.cc/embed/dlBQdytGVnBONks3b1JuNDZNT1JmaXVkSjU0SGlxUThXblJFVWxVcjl1UT0" frameborder="0" webkitallowfullscreen mozallowfullscreen allowfullscreen></iframe>