ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

C++多线程实现绑定CPU的方法详解

C++多线程实现绑定CPU的方法详解 前言先把话说清楚C 标准里没有任何 CPU 亲和性CPU affinity的 API。 标准库的std::thread不提供把这个线程绑到第 N 个核的能力std::thread::hardware_concurrency()也只能告诉你逻辑处理器大概有几个。 所以用 C 绑定 CPU这个说法的准确含义是用std::thread管理线程 再通过std::thread::native_handle()拿到平台原生句柄 调用操作系统的亲和性接口完成绑定。离开操作系统 API这件事在纯标准 C 里做不到。典型误解有两个。第一以为存在某个std::this_thread::set_affinity()——不存在的。 第二以为绑核就是独占这个核——不是。亲和性只是限定调度器的候选集合 那个核上仍然可以跑别的线程除非你用 cgroup 的 cpuset 或内核的isolcpus把它隔离出来。本文以 C17 为基准分别给出 Linuxglibc与 Windows 两条路径 并重点讲清一个容易被忽略的时序问题std::thread构造完成时线程可能已经开始跑了 此时再去设置亲和性存在竞态。一、为什么需要绑定以及它的代价绑定的收益来自硬件层面的两个事实缓存局部性线程被调度器在不同 CPU 之间迁移时原先在某个核的私有L1/L2 缓存里预热的数据就凉了迁移后要重新从更远的层级取。NUMA 亲和性在多路服务器上跨 NUMA 节点访问内存的延迟高于本地节点。把线程钉在某个节点附近的核上有助于让它更多地命中本地内存。此外实时和低延迟场景希望减少调度抖动性能测量场景希望降低噪声 这些都属于用可控性换取灵活性的考虑。反过来绑定是有代价的写之前要想清楚绑得太死调度器没法把线程挪到空闲核上。如果被绑定的核上还有别的忙线程你的线程只能干等而旁边有核闲着。一个核上绑多个忙线程会退化成线程间的相互抢占比不绑定还糟。二、Linuxpthread_setaffinity_np 与 sched_setaffinityglibc 提供两个亲和性接口int pthread_setaffinity_np(pthread_t thread, size_t cpusetsize, const cpu_set_t *cpuset); int pthread_getaffinity_np(pthread_t thread, size_t cpusetsize, cpu_set_t *cpuset); int sched_setaffinity(pid_t pid, size_t cpusetsize, const cpu_set_t *mask);几个必须记住的细节这两个*_np函数是GNU 扩展non-portable所以带 np 后缀使用前必须在任何#include之前#define _GNU_SOURCE且需要-pthread链接选项。它们返回错误码本身成功返回 0不设置errno。把返回值当 errno 用是常见的错。cpu_set_t是一张位图。操作宏有CPU_ZERO(set)、CPU_SET(cpu, set)、CPU_CLR(cpu, set)、CPU_ISSET(cpu, set)、CPU_COUNT(set) 都在sched.h里。这些宏不做边界检查越界写会破坏栈或相邻内存——是未定义行为。cpusetsize必须传sizeof(cpu_set_t)或者你用CPU_ALLOC动态分配得到的字节数。sched_setaffinity的pid传0 表示调用线程自己。这个重载的好处是在线程函数里调用它没有任何竞态不需要先拿到自己的线程 ID。查询当前线程实际跑在哪个核上可以用 glibc 的sched_getcpu()。三、WindowsSetThreadAffinityMask 与处理器组Win32 侧的对应接口DWORD_PTR SetThreadAffinityMask(HANDLE hThread, DWORD_PTR dwThreadAffinityMask); BOOL SetThreadGroupAffinity(HANDLE hThread, const GROUP_AFFINITY *GroupAffinity, PGROUP_AFFINITY PreviousGroupAffinity); BOOL GetThreadGroupAffinity(HANDLE hThread, PGROUP_AFFINITY GroupAffinity);要点SetThreadAffinityMask成功时返回原来的掩码失败返回0。Windows 没有单独的 GetThreadAffinityMask要查询就先按原值设置一次 返回值就是当前掩码。掩码是位图第 n 位为 1 表示允许跑在第 n 个逻辑处理器上。GetCurrentThread()返回的是伪句柄pseudo handle它永久有效不要对它调用CloseHandle。当逻辑处理器数量超过 64 时Windows 引入处理器组processor group概念一个掩码只能覆盖一个组内最多 64 个处理器。跨组要用SetThreadGroupAffinity 用GetActiveProcessorGroupCount()与GetActiveProcessorCount(group)查询规模GetActiveProcessorMask()拿到当前组的掩码。用GetLastError()取失败原因。四、和 std::thread 拼接竞态与三种方案std::thread::native_handle()返回的就是原生句柄在 Linux 上libstdc 与 libc 把native_handle_type定义为pthread_t这是实现细节不是标准规定可用static_assert(std::is_samestd::thread::native_handle_type, pthread_t::value, );自行确认。在 MSVC STL 上native_handle_type被定义为void*实际值就是线程的HANDLE同样是实现细节以本机thread头文件为准。于是有了第一种写法父线程拿到句柄后设置。std::thread t(worker); pthread_setaffinity_np(t.native_handle(), sizeof(set), set); // 有竞态问题在于std::thread的构造函数返回时新线程可能已经在跑了。 在这行代码执行之前被绑定的线程可能已经在任意核上工作了很久。 对只是希望它尽量在某个核上的普通场景这通常无所谓 但对从第一条指令起就必须在指定核上的实时场景这就是个真问题。三种解法线程自绑推荐无竞态在线程函数开头调用pthread_setaffinity_np(pthread_self(), ...)或SetThreadAffinityMask(GetCurrentThread(), ...)。 或者用sched_setaffinity的 pid 传 0 的写法。加一道门父线程先设置native_handle的亲和性再用互斥量加条件变量把子线程放行。绕开 std::thread直接用pthread_create或_beginthreadex创建线程在创建前后自己控制一切。可移植性最差。顺带提醒一个相关话题绑定 CPU 常常和让工作线程长期钉在某个核上跑一起出现 而std::thread没有标准的取消机制。C20 的std::jthread也没有抢占式取消 它只是提供了一个协作式的std::stop_token。要让线程停下来只能自己准备一个std::atomicbool标志由线程周期性检查std::atomicbool stop_flag{false}; // 需要 atomic std::thread worker([] { while (!stop_flag.load(std::memory_order_relaxed)) { do_one_unit_of_work(); } }); // ... 结束阶段 stop_flag.store(true, std::memory_order_relaxed); // 只是通知停止不传数据 worker.join();注意两点volatile不能替代std::atomic它既不保证原子性也不建立 happens-before 关系另外如果这个标志还要用来发布主线程写好的数据 就必须用release/acquire配对relaxed只适合纯粹的停止信号。五、完整示例Linux 版// 适用Linux glibc GCC 13 / Clang 17 // 编译g -stdc17 -O2 -Wall -Wextra -pthread affinity_linux.cpp -o affinity_linux #define _GNU_SOURCE // 必须在包含 sched.h / pthread.h 之前定义 #include pthread.h #include sched.h #include condition_variable #include cstdio #include cstring #include mutex #include thread namespace { // 把「当前线程」绑定到编号为 cpu 的逻辑处理器上无竞态 bool bind_this_thread(int cpu) { if (cpu 0 || cpu CPU_SETSIZE) { // CPU_SET 宏不做边界检查自己先挡 return false; } cpu_set_t set; CPU_ZERO(set); CPU_SET(cpu, set); const int rc ::pthread_setaffinity_np(::pthread_self(), sizeof(set), set); if (rc ! 0) { // 返回错误码本身不设置 errno std::fprintf(stderr, pthread_setaffinity_np 失败: %s\n, std::strerror(rc)); return false; } return true; } } // namespace int main() { if (std::thread::hardware_concurrency() 0) { std::fprintf(stderr, 无法确定逻辑处理器数量\n); return 1; } const int target 0; // 目标逻辑处理器编号 std::mutex m; std::condition_variable cv; bool ready false; std::thread worker([] { { // 先在这里等确保父线程已经完成对它 native_handle 的设置 std::unique_lockstd::mutex lk(m); cv.wait(lk, [] { return ready; }); } // 路径一线程自绑重复设置同一掩码是幂等的 if (!bind_this_thread(target)) { std::fprintf(stderr, 线程自绑失败\n); } std::printf(worker 运行在 CPU %d\n, ::sched_getcpu()); }); // 路径二父线程先用 native_handle 设置亲和性再放行 cpu_set_t set; CPU_ZERO(set); CPU_SET(target, set); const int rc ::pthread_setaffinity_np(worker.native_handle(), sizeof(set), set); if (rc ! 0) { std::fprintf(stderr, 设置子线程亲和性失败: %s\n, std::strerror(rc)); } { std::lock_guardstd::mutex lk(m); ready true; } cv.notify_one(); worker.join(); return 0; }Windows 版// 适用Windows MSVC 19.3xMinGW-w64 亦可无需额外链接库 // 编译cl /std:c17 /EHsc /W4 affinity_win.cpp #define WIN32_LEAN_AND_MEAN #include windows.h #include condition_variable #include cstdio #include mutex #include thread namespace { bool bind_current_thread(unsigned cpu) { if (cpu 64) { // 一个处理器组内最多 64 个逻辑处理器 return false; // 跨组请改用 SetThreadGroupAffinity } const DWORD_PTR mask DWORD_PTR(1) cpu; // 成功返回原来的掩码失败返回 0 return ::SetThreadAffinityMask(::GetCurrentThread(), mask) ! 0; } } // namespace int main() { const unsigned target 0; std::mutex m; std::condition_variable cv; bool ready false; std::thread worker([] { { std::unique_lockstd::mutex lk(m); cv.wait(lk, [] { return ready; }); } // 路径一线程自绑无竞态 if (!bind_current_thread(target)) { std::fprintf(stderr, 线程自绑失败\n); } std::printf(worker 已绑定到 CPU %u\n, target); }); // 路径二父线程通过 native_handle 再确认一次幂等用于演示句柄的取法 // MSVC STL 中 native_handle_type 是 void*值就是线程 HANDLE HANDLE h static_castHANDLE(worker.native_handle()); const DWORD_PTR mask DWORD_PTR(1) target; if (::SetThreadAffinityMask(h, mask) 0) { std::fprintf(stderr, SetThreadAffinityMask 失败, 错误码 %lu\n, static_castunsigned long(::GetLastError())); } { std::lock_guardstd::mutex lk(m); ready true; } cv.notify_one(); worker.join(); return 0; }两个示例都用了同一套门闩结构。要验证绑定是否生效Linux 上用sched_getcpu()或外部命令taskset -pWindows 上用SetThreadAffinityMask传入原掩码再看它的返回值。常见坑点_GNU_SOURCE定义得太晚❌ 先#include sched.h再#define _GNU_SOURCE——pthread_setaffinity_np声明不出来报隐式声明或未定义标识符。✅ 把#define _GNU_SOURCE放在文件最顶部、所有#include之前。把pthread_setaffinity_np的返回值当 errno❌if (pthread_setaffinity_np(...) -1) { perror(bind); }—— 它成功返回 0、 失败返回错误码本身从不返回 -1也不设置errno。✅const int rc pthread_setaffinity_np(...); if (rc ! 0) fprintf(stderr, %s\n, strerror(rc));父线程设置子线程亲和性时忽略竞态❌std::thread t(f); pthread_setaffinity_np(t.native_handle(), ...);—— 构造函数返回时子线程可能已经执行了很久绑定只对之后生效。✅ 让线程自己绑pthread_self()或GetCurrentThread() 或者用互斥量加条件变量建一道门等父线程设置完再放行。CPU_SET越界❌CPU_SET(2048, set);—— 宏只做位运算不做检查写穿了cpu_set_t的边界这是未定义行为症状可能是栈被破坏。✅ 先判断编号在[0, CPU_SETSIZE)内核数很多时用CPU_ALLOC/CPU_FREE动态分配位图并把实际字节数传给cpusetsize。Windows 上对GetCurrentThread()的伪句柄调用CloseHandle❌HANDLE h GetCurrentThread(); ... CloseHandle(h);—— 伪句柄不是真实句柄。✅ 伪句柄永远有效、永不需要关闭std::thread::native_handle()拿到的线程句柄 由标准库在线程对象析构/join/detach 时负责关闭你也不要手动CloseHandle。在 Windows 上把掩码移位到 64 位以上❌DWORD_PTR(1) 64—— 移位量大于等于位宽是未定义行为 而且 64 位掩码在处理器组模型里本身就不表示第 64 个核。✅ 守卫cpu 64超出时改用SetThreadGroupAffinity指定组号和组内掩码。想用std::thread直接取消一个已绑核的线程❌t.kill()、t.cancel()、t.terminate()—— 标准库没有这些接口 编译都过不了Windows 上的TerminateThread是原生 API强杀会泄漏资源、 留下锁未释放的线程几乎总是坏主意。✅ 用std::atomicbool做协作式退出标志线程周期性检查 主线程置位后join()等待收尾。把hardware_concurrency()的结果当可用核数❌const unsigned n std::thread::hardware_concurrency();然后按 n 均分线程 —— 它可能返回 0标准允许而且不看cgroup / 亲和性掩码等限制 在容器里它常常是宿主机核数而不是配额。✅ 判 0时退回 1需要真实可用核数时Linux 读 cgroup/sched_getaffinity的结果Windows 用GetActiveProcessorCount。总结平台 / 层面接口关键注意点C 标准无标准库不提供亲和性只能靠native_handle()借道Linuxglibcpthread_setaffinity_np、sched_setaffinityGNU 扩展需_GNU_SOURCE与-pthread返回错误码而非 errnoLinux 自绑sched_setaffinity(0, ...)无竞态最省事WindowsSetThreadAffinityMask返回原掩码GetCurrentThread()是伪句柄Windows多组SetThreadGroupAffinity超过 64 个逻辑处理器时需要竞态处理线程自绑 或 条件变量门闩std::thread构造返回时线程可能已在运行一句话总结标准 C 只管创建和汇合线程绑核是操作系统的事。 实践的推荐顺序是——先让线程自己绑最可靠必须由外部指定时再用native_handle()配合门闩同时牢记std::thread没有取消机制 停止线程永远是协作式的。
返回列表