Skip to content

Getting Started with OpenMP

In shared-memory MIMD (Multiple Instruction, Multiple Data) systems, all processors or CPU cores share access to a globally accessible main memory. Each thread executing within a process can potentially read from and write to any location in this shared address space.

flowchart TD
  subgraph SharedSystem["Figure 5.1: Shared-Memory Architecture"]
      direction TB
      CPU0["CPU / Core 0"] <==> BUS["Interconnect / Bus"]
      CPU1["CPU / Core 1"] <==> BUS
      CPU2["CPU / Core 2"] <==> BUS
      CPUn["CPU / Core p-1"] <==> BUS
      BUS <==> MEM["Globally Accessible Main Memory"]
  end

To program shared-memory machines, developers typically choose between two standard C APIs: Pthreads (POSIX Threads) and OpenMP (Open Multi-Processing).

FeaturePOSIX Threads (Pthreads)OpenMP
API TypeLibrary of function callsCompiler directives (#pragma) + runtime library
Abstraction LevelLow-level (explicit thread management)High-level (declarative worksharing)
Thread ManagementManual creation, attribute setup, join loopsAutomatic thread pool management by runtime
Code ModificationRequires extensive restructuring of serial codeEnables incremental parallelization of serial code
Compiler RequirementAny standard C compiler with POSIX libraryRequires compiler support (e.g., -fopenmp in GCC/Clang)

OpenMP was designed specifically to allow programmers to incrementally parallelize existing serial applications: developers can identify computational bottlenecks (such as large loops) and insert compiler directives with minimal changes to the original source code.


5.1 Directives-Based Parallelism and Pragmas

Section titled “5.1 Directives-Based Parallelism and Pragmas”

OpenMP relies primarily on compiler directives known as pragmas. In C and C++, pragmas instruct the compiler to perform platform-specific actions outside the standard language grammar:

#pragma omp <directive-name> [clause ...]
  • Pragmas begin with #pragma omp starting in column 1.
  • By default, a pragma occupies a single line. Long directives can be split across multiple lines by escaping the newline with a trailing backslash (\):
#pragma omp parallel num_threads(thread_count) \
default(none) shared(a, b) private(i)

5.2 A First OpenMP Program: Parallel Greetings

Section titled “5.2 A First OpenMP Program: Parallel Greetings”

Below is the complete C source code for Program 5.1 (omp_hello.c), demonstrating basic thread creation and identification:

/* Program 5.1: A "hello, world" program that uses OpenMP */
#include <stdio.h>
#include <stdlib.h>
#include <omp.h>
void Hello(void); /* Prototype for thread function */
int main(int argc, char* argv[]) {
/* Read requested thread count from command line */
int thread_count = strtol(argv[1], NULL, 10);
/* Fork a team of threads to execute Hello() */
#pragma omp parallel num_threads(thread_count)
Hello();
return 0;
} /* main */
void Hello(void) {
int my_rank = omp_get_thread_num();
int thread_count = omp_get_num_threads();
printf("Hello from thread %d of %d\n", my_rank, thread_count);
} /* Hello */

5.2.1 Compiling and Running OpenMP Programs

Section titled “5.2.1 Compiling and Running OpenMP Programs”

To compile OpenMP code using GCC or Clang, provide the -fopenmp flag:

Terminal window
$ gcc -g -Wall -fopenmp -o omp_hello omp_hello.c

The program is executed by supplying the desired number of threads as a command-line argument:

Terminal window
$ ./omp_hello 4
Hello from thread 0 of 4
Hello from thread 1 of 4
Hello from thread 2 of 4
Hello from thread 3 of 4

Because all threads execute concurrently and compete for access to standard output (stdout), their print statements are non-deterministic. Consecutive executions may produce different permutations:

Run 1: Run 2:
Hello from thread 1 of 4 Hello from thread 3 of 4
Hello from thread 2 of 4 Hello from thread 0 of 4
Hello from thread 0 of 4 Hello from thread 1 of 4
Hello from thread 3 of 4 Hello from thread 2 of 4

OpenMP programs execute according to the Fork-Join execution model:

flowchart LR
  subgraph ForkJoin["Figure 5.2: Fork-Join Threading Model"]
      direction LR
      P1["Initial Master Thread"] --> FORK{"#pragma omp parallel
(Fork)"}
      FORK --> T0["Master Thread (Rank 0)"]
      FORK --> T1["Child Thread (Rank 1)"]
      FORK --> T2["Child Thread (Rank 2)"]
      FORK --> T3["Child Thread (Rank 3)"]
      T0 --> JOIN{"Implicit Barrier
(Join)"}
      T1 --> JOIN
      T2 --> JOIN
      T3 --> JOIN
      JOIN --> P2["Master Thread Continues"]
  end
  1. Sequential Start: The application begins execution as a single-threaded process run by the initial thread (called the master thread or parent thread).
  2. Fork: When the master thread encounters a #pragma omp parallel directive, it creates a team of threads. The master thread retains rank 0, and thread_count - 1 additional child threads are spawned.
  3. Parallel Execution: Every thread in the team concurrently executes the structured block immediately following the directive.
  4. Join & Implicit Barrier: At the end of the parallel block, an implicit barrier forces all threads to wait until the entire team finishes. Child threads terminate, and only the master thread continues sequentially beyond the block.

The OpenMP runtime header <omp.h> provides environment inquiry functions:

/* Returns the unique integer ID (rank) of the calling thread: 0 <= rank < thread_count */
int omp_get_thread_num(void);
/* Returns the total number of threads in the current active team */
int omp_get_num_threads(void);

Each thread has its own private call stack. Therefore, variables declared inside functions called within a parallel region (such as my_rank and thread_count inside Hello()) reside on the thread’s private stack and are automatically private to that thread.


5.2.4 Error Checking and Portability with _OPENMP

Section titled “5.2.4 Error Checking and Portability with _OPENMP”

In production software, robust error checking should verify command-line arguments and handle environments where OpenMP is unavailable.

OpenMP-compliant compilers define the preprocessor macro _OPENMP. We can use conditional compilation (#ifdef _OPENMP) to ensure our code compiles cleanly even on legacy compilers:

#ifdef _OPENMP
#include <omp.h>
#endif
int main(int argc, char* argv[]) {
int thread_count;
if (argc != 2) {
fprintf(stderr, "usage: %s <thread_count>\n", argv[0]);
exit(1);
}
thread_count = strtol(argv[1], NULL, 10);
if (thread_count <= 0) {
fprintf(stderr, "thread count must be > 0\n");
exit(1);
}
# pragma omp parallel num_threads(thread_count)
Hello();
return 0;
}
void Hello(void) {
# ifdef _OPENMP
int my_rank = omp_get_thread_num();
int thread_count = omp_get_num_threads();
# else
int my_rank = 0;
int thread_count = 1;
# endif
printf("Hello from thread %d of %d\n", my_rank, thread_count);
}

If compiled without OpenMP support (gcc omp_hello.c), the compiler safely ignores #pragma omp, sets my_rank = 0 and thread_count = 1, and runs the program sequentially without errors.