Parallel Computing Course
Learn the foundations of parallel computing — from why we need ever-increasing performance to writing your own parallel programs — with interactive code editors and visual explanations.
Chapter 1: Introduction to Parallel Computing
Section titled “Chapter 1: Introduction to Parallel Computing”1. The Need for Ever-Increasing PerformanceWhy parallel computing? The growing demand for performance, the physical limits of monolithic processors (Power Wall) and the rise of multicore systems.
2. Writing Parallel ProgramsWork-partitioning paradigms (data parallelism and task parallelism), coordination, synchronization, load balancing and SIMD/MIMD architectures.
Chapter 2: Distributed Memory Programming with MPI
Section titled “Chapter 2: Distributed Memory Programming with MPI”1. Getting Started with MPIFundamentals of distributed memory, message passing, SPMD execution, and point-to-point communication with MPI_Send and MPI_Recv.
2. Trapezoidal Rule & Distributed I/OParallelizing numerical integration using Foster's methodology, distinguishing local vs. global variables, and handling console I/O.
3. Collective CommunicationOptimized tree reductions, MPI_Reduce, MPI_Bcast, vector distributions, scatter/gather operations, and parallel matrix-vector multiplication.
4. Derived Datatypes & PerformanceConsolidating heterogeneous data with MPI structs, high-resolution wall-clock timing, speedup, efficiency, and scalability analysis.
5. Parallel Sorting & Communication SafetyParallel odd-even transposition sort, deadlock avoidance, communication scheduling, and safe message exchange with MPI_Sendrecv.
Chapter 3: Shared-Memory Programming with OpenMP
Section titled “Chapter 3: Shared-Memory Programming with OpenMP”1. Getting Started with OpenMPShared-memory MIMD systems, pragma directives, the fork-join threading model, thread ID inquiry, and compiler portability with _OPENMP.
2. Trapezoidal Rule & ReductionsNumerical integration in shared memory, race condition mechanics, mutual exclusion with critical, variable scoping, and reduction clauses.
3. Parallel for & Loop SchedulingCanonical loop restrictions, loop-carried data dependences, calculating Pi, default(none) scoping, and static/dynamic/guided scheduling.
4. Synchronization, Locks & QueuesProducer-consumer queue pattern, explicit barriers, hardware-accelerated atomic operations, named critical sections, and OpenMP locks.
5. False Sharing, Tasking & SafetyCache coherence and false sharing, thread reuse in odd-even sorting, dynamic OpenMP 3.0 Tasking API, and reentrant thread-safety.
Chapter 4: GPU Programming with CUDA
Section titled “Chapter 4: GPU Programming with CUDA”1. GPU Architecture & CUDA BasicsMassively parallel GPU computing, SIMD vs SIMT execution, Streaming Multiprocessors, host/device model, and launching kernels with thread hierarchies.
2. Vector Addition & Memory ManagementVector addition on GPUs, global thread indexing, Unified Memory with cudaMallocManaged, and explicit device memory transfers with cudaMemcpy.
3. Trapezoidal Rule & AtomicsNumerical integration on thousands of GPU threads, race conditions on global memory accumulators, hardware atomicAdd(), and serialization bottlenecks.
4. Memory Hierarchy & Warp ShufflesTree-structured reductions, GPU memory hierarchy (registers, shared, global), warp execution, and ultrafast register warp shuffles (__shfl_down_sync).
5. Multi-Warp Reductions & Bitonic SortShared memory bank conflict avoidance, barrier synchronization with __syncthreads(), and high-performance parallel Bitonic Sorting across GPU grids.
Chapter 5: Parallel Program Development
Section titled “Chapter 5: Parallel Program Development”1. n-Body Problem & Shared-MemoryGravitational n-body physics, basic vs reduced force calculations, Foster's methodology, and two-phase private force accumulation in OpenMP & Pthreads.
2. Distributed n-Body & Ring PassDistributed-memory n-body simulation in MPI, in-place allgather optimization, the ring pass pipeline topology, and memory scalability comparisons.
3. GPU n-Body & Shared Memory TilingFine-grained CUDA kernel design, host-coordinated grid barriers, GPU memory limits, and programmer-managed cache tiling using on-chip shared memory.
4. Sample Sort & Shared-MemoryBucket sort generalization, sample and splitter selection, exclusive prefix sums, count matrices, and high-performance OpenMP/Pthreads sorting.
5. Distributed & GPU Sort, API StrategyButterfly all-to-all exchanges in MPI, multi-block GPU sample sort in CUDA, and a comprehensive decision framework for parallel API selection.