Getting Started with OpenMP
In shared-memory MIMD (Multiple Instruction, Multiple Data) systems, all processors or CPU cores share access to a globally accessible main memory. Each thread executing within a process can potentially read from and write to any location in this shared address space.
flowchart TD
subgraph SharedSystem["Figure 5.1: Shared-Memory Architecture"]
direction TB
CPU0["CPU / Core 0"] <==> BUS["Interconnect / Bus"]
CPU1["CPU / Core 1"] <==> BUS
CPU2["CPU / Core 2"] <==> BUS
CPUn["CPU / Core p-1"] <==> BUS
BUS <==> MEM["Globally Accessible Main Memory"]
endTo program shared-memory machines, developers typically choose between two standard C APIs: Pthreads (POSIX Threads) and OpenMP (Open Multi-Processing).
| Feature | POSIX Threads (Pthreads) | OpenMP |
|---|---|---|
| API Type | Library of function calls | Compiler directives (#pragma) + runtime library |
| Abstraction Level | Low-level (explicit thread management) | High-level (declarative worksharing) |
| Thread Management | Manual creation, attribute setup, join loops | Automatic thread pool management by runtime |
| Code Modification | Requires extensive restructuring of serial code | Enables incremental parallelization of serial code |
| Compiler Requirement | Any standard C compiler with POSIX library | Requires compiler support (e.g., -fopenmp in GCC/Clang) |
OpenMP was designed specifically to allow programmers to incrementally parallelize existing serial applications: developers can identify computational bottlenecks (such as large loops) and insert compiler directives with minimal changes to the original source code.
5.1 Directives-Based Parallelism and Pragmas
Section titled “5.1 Directives-Based Parallelism and Pragmas”OpenMP relies primarily on compiler directives known as pragmas. In C and C++, pragmas instruct the compiler to perform platform-specific actions outside the standard language grammar:
#pragma omp <directive-name> [clause ...]- Pragmas begin with
#pragma ompstarting in column 1. - By default, a pragma occupies a single line. Long directives can be split across multiple lines by escaping the newline with a trailing backslash (
\):
#pragma omp parallel num_threads(thread_count) \ default(none) shared(a, b) private(i)5.2 A First OpenMP Program: Parallel Greetings
Section titled “5.2 A First OpenMP Program: Parallel Greetings”Below is the complete C source code for Program 5.1 (omp_hello.c), demonstrating basic thread creation and identification:
/* Program 5.1: A "hello, world" program that uses OpenMP */#include <stdio.h>#include <stdlib.h>#include <omp.h>
void Hello(void); /* Prototype for thread function */
int main(int argc, char* argv[]) { /* Read requested thread count from command line */ int thread_count = strtol(argv[1], NULL, 10);
/* Fork a team of threads to execute Hello() */ #pragma omp parallel num_threads(thread_count) Hello();
return 0;} /* main */
void Hello(void) { int my_rank = omp_get_thread_num(); int thread_count = omp_get_num_threads();
printf("Hello from thread %d of %d\n", my_rank, thread_count);} /* Hello */5.2.1 Compiling and Running OpenMP Programs
Section titled “5.2.1 Compiling and Running OpenMP Programs”To compile OpenMP code using GCC or Clang, provide the -fopenmp flag:
$ gcc -g -Wall -fopenmp -o omp_hello omp_hello.cThe program is executed by supplying the desired number of threads as a command-line argument:
$ ./omp_hello 4Hello from thread 0 of 4Hello from thread 1 of 4Hello from thread 2 of 4Hello from thread 3 of 4Because all threads execute concurrently and compete for access to standard output (stdout), their print statements are non-deterministic. Consecutive executions may produce different permutations:
Run 1: Run 2:Hello from thread 1 of 4 Hello from thread 3 of 4Hello from thread 2 of 4 Hello from thread 0 of 4Hello from thread 0 of 4 Hello from thread 1 of 4Hello from thread 3 of 4 Hello from thread 2 of 45.2.2 The Fork-Join Model
Section titled “5.2.2 The Fork-Join Model”OpenMP programs execute according to the Fork-Join execution model:
flowchart LR
subgraph ForkJoin["Figure 5.2: Fork-Join Threading Model"]
direction LR
P1["Initial Master Thread"] --> FORK{"#pragma omp parallel
(Fork)"}
FORK --> T0["Master Thread (Rank 0)"]
FORK --> T1["Child Thread (Rank 1)"]
FORK --> T2["Child Thread (Rank 2)"]
FORK --> T3["Child Thread (Rank 3)"]
T0 --> JOIN{"Implicit Barrier
(Join)"}
T1 --> JOIN
T2 --> JOIN
T3 --> JOIN
JOIN --> P2["Master Thread Continues"]
end- Sequential Start: The application begins execution as a single-threaded process run by the initial thread (called the master thread or parent thread).
- Fork: When the master thread encounters a
#pragma omp paralleldirective, it creates a team of threads. The master thread retains rank0, andthread_count - 1additional child threads are spawned. - Parallel Execution: Every thread in the team concurrently executes the structured block immediately following the directive.
- Join & Implicit Barrier: At the end of the parallel block, an implicit barrier forces all threads to wait until the entire team finishes. Child threads terminate, and only the master thread continues sequentially beyond the block.
5.2.3 Core OpenMP Runtime Functions
Section titled “5.2.3 Core OpenMP Runtime Functions”The OpenMP runtime header <omp.h> provides environment inquiry functions:
/* Returns the unique integer ID (rank) of the calling thread: 0 <= rank < thread_count */int omp_get_thread_num(void);
/* Returns the total number of threads in the current active team */int omp_get_num_threads(void);Each thread has its own private call stack. Therefore, variables declared inside functions called within a parallel region (such as my_rank and thread_count inside Hello()) reside on the thread’s private stack and are automatically private to that thread.
5.2.4 Error Checking and Portability with _OPENMP
Section titled “5.2.4 Error Checking and Portability with _OPENMP”In production software, robust error checking should verify command-line arguments and handle environments where OpenMP is unavailable.
OpenMP-compliant compilers define the preprocessor macro _OPENMP. We can use conditional compilation (#ifdef _OPENMP) to ensure our code compiles cleanly even on legacy compilers:
#ifdef _OPENMP #include <omp.h>#endif
int main(int argc, char* argv[]) { int thread_count;
if (argc != 2) { fprintf(stderr, "usage: %s <thread_count>\n", argv[0]); exit(1); }
thread_count = strtol(argv[1], NULL, 10); if (thread_count <= 0) { fprintf(stderr, "thread count must be > 0\n"); exit(1); }
# pragma omp parallel num_threads(thread_count) Hello();
return 0;}
void Hello(void) {# ifdef _OPENMP int my_rank = omp_get_thread_num(); int thread_count = omp_get_num_threads();# else int my_rank = 0; int thread_count = 1;# endif
printf("Hello from thread %d of %d\n", my_rank, thread_count);}If compiled without OpenMP support (gcc omp_hello.c), the compiler safely ignores #pragma omp, sets my_rank = 0 and thread_count = 1, and runs the program sequentially without errors.