
简介本资源是一份面向计算机视觉初学者与YOLOv5模型实践者的行人检测专用数据集适用于目标检测算法训练、模型调优及课程实验等场景。数据集共包含2000张真实场景行人图像JPG格式配套2095个YOLOv5标准标签文件TXT与2095个PASCAL VOC格式标注文件XML便于在不同框架间灵活迁移另有少量辅助文件总计6287个文件压缩包大小为180.61MB结构规整、开箱即用。目前已有597人学习下载体现了社区对轻量级行人检测基准数据的持续需求。用户可直接划分训练/验证/测试集快速启动YOLOv5s或YOLOv5m模型训练所有图片经人工筛选与标注校验覆盖多种光照、角度与遮挡条件适合作为入门级检测任务的数据基线也便于拓展为多类别行人属性识别的预研基础。1. 2000张行人数据集YOLOv5格式不是拿来就能训而是得先看清它到底“长什么样”你下载了一个叫行人数据集2000张-yolov5格式数据集.zip的压缩包解压后看到images/和labels/两个文件夹.jpg图片和对应.txt标签一一配对——第一反应是“终于能直接喂给 YOLOv5 训了”。但现实很快打脸训练 loss 疯狂震荡、mAP 停在 0.1 不动、验证时框满天飞却全打偏……问题不在模型而在这个“2000张”本身——它大概率是未经清洗的原始采集片段含大量遮挡、小目标、模糊帧、重复样本且标签坐标可能越界、归一化错误、类别ID写错。这不是一个开箱即用的数据集而是一份需要你亲手“验伤、清创、缝合”的临床标本。本文不讲 YOLOv5 安装或训练命令只聚焦于如何用最小成本确认这份行人数据集2000张-yolov5格式数据集.zip是否真实可用若不可用怎样在不重采、不重标前提下用代码人工交叉验证快速修复适合刚跑通官方 demo、正准备训自己场景的工程师也适合被甲方甩来一堆“已标注好”的数据、却不敢贸然投入训练周期的算法负责人。2. 解构 ZIP 包先看结构再动手避免“解压即失败”拿到 ZIP 文件第一件事不是unzip而是用file和zipinfo快速验明正身。很多所谓“YOLOv5格式”数据集实际是用脚本批量转换时出错导致labels/里.txt文件名与images/中.jpg不严格一致大小写、空格、中文编码或labels/下混入.DS_Store、.gitignore等隐藏文件——这些都会让torch.utils.data.Dataset在__getitem__时直接报FileNotFoundError且错误堆栈藏在DataLoader底层极难定位。2.1 检查 ZIP 内部结构与文件一致性# 查看 ZIP 内容结构不解压 zipinfo 行人数据集2000张-yolov5格式数据集.zip | head -30 # 提取文件列表检查 images/ 和 labels/ 是否成对存在 unzip -l 行人数据集2000张-yolov5格式数据集.zip | grep -E (images/|labels/) | awk {print $4} | sort file_list.txt提示重点观察三处①images/和labels/是否同级目录② 所有.jpg是否都在images/下而非images/train/或JPEGImages/③labels/中是否只有.txt且无.xml、.json混入。2.2 验证 YOLOv5 格式核心约束坐标归一化 类别ID合法性YOLOv5 要求每张图的.txt标签文件中每行格式为class_id x_center y_center width height其中x_center,y_center,width,height必须是0~1 之间的浮点数归一化到图像宽高且class_id必须为0行人唯一类别。常见错误包括坐标未归一化直接用了像素值x_center或y_center小于 0 或大于 1width或height为 0 或负数class_id写成1或person字符串。我们写一个轻量校验脚本不依赖cv2或PIL纯用 Python 标准库读取尺寸并验证# validate_yolo_format.py import os import glob from pathlib import Path def check_yolo_labels(img_dir, label_dir): img_files list(Path(img_dir).glob(*.jpg)) list(Path(img_dir).glob(*.jpeg)) label_files list(Path(label_dir).glob(*.txt)) # 构建文件名映射忽略扩展名 img_stems {img.stem for img in img_files} label_stems {lbl.stem for lbl in label_files} missing_labels img_stems - label_stems missing_images label_stems - img_stems print(f[INFO] 缺失 label 的图片: {len(missing_labels)} 张) print(f[INFO] 缺失 image 的 label: {len(missing_images)} 张) # 逐个检查 label 内容 invalid_lines [] for lbl_path in label_files: try: with open(lbl_path, r) as f: lines f.readlines() for i, line in enumerate(lines): parts line.strip().split() if len(parts) ! 5: invalid_lines.append(f{lbl_path.name}:{i1} → 字段数≠5 ({len(parts)})) continue try: cls_id int(parts[0]) x, y, w, h map(float, parts[1:]) if cls_id ! 0: invalid_lines.append(f{lbl_path.name}:{i1} → class_id{cls_id} (应为0)) if not (0 x 1 and 0 y 1 and 0 w 1 and 0 h 1): invalid_lines.append(f{lbl_path.name}:{i1} → 坐标越界 x{x:.3f} y{y:.3f} w{w:.3f} h{h:.3f}) except ValueError: invalid_lines.append(f{lbl_path.name}:{i1} → 含非数字字符) except Exception as e: invalid_lines.append(f{lbl_path.name} → 读取异常: {e}) print(f[INFO] 共发现 {len(invalid_lines)} 处格式错误:) for err in invalid_lines[:10]: # 只打印前10条避免刷屏 print(f • {err}) if len(invalid_lines) 10: print(f ... 还有 {len(invalid_lines)-10} 条未显示) if __name__ __main__: # 替换为你解压后的路径 IMG_DIR 行人数据集2000张-yolov5格式数据集/images LABEL_DIR 行人数据集2000张-yolov5格式数据集/labels check_yolo_labels(IMG_DIR, LABEL_DIR)运行后关键看三类输出缺失 label 的图片说明有图没标需补标或剔除缺失 image 的 label说明有标没图属脏数据直接删坐标越界最常见问题根源常是标注工具导出 bug 或图像 resize 后未同步更新 label。参数说明脚本默认class_id0若你后续要扩展其他类别如“骑车人”此处需改为cls_id in [0,1]并同步修改data.yaml0 w 1中的0 w是硬性要求——YOLOv5 会因w0报ZeroDivisionError。3. 数据质量诊断用可视化统计揪出“伪2000张”里的水分2000 张听起来不少但若 80% 是同一摄像头下的连续帧时间冗余、30% 目标尺寸小于 20×20 像素小目标噪声、50% 图像分辨率低于 640×480低质图像实际有效样本可能不足 500 张。必须量化评估而非凭感觉。3.1 统计图像基础属性尺寸、比例、亮度分布# stats_image_quality.py import cv2 import numpy as np from pathlib import Path import matplotlib.pyplot as plt def analyze_images(img_dir): img_paths list(Path(img_dir).glob(*.jpg)) list(Path(img_dir).glob(*.jpeg)) widths, heights, ratios, means [], [], [], [] for p in img_paths[:500]: # 先抽样500张避免卡死 try: img cv2.imread(str(p)) if img is None: continue h, w img.shape[:2] widths.append(w) heights.append(h) ratios.append(w/h) means.append(np.mean(img)) except: continue # 绘制分布直方图 fig, axes plt.subplots(2, 2, figsize(12, 8)) axes[0,0].hist(widths, bins50); axes[0,0].set_title(Width Distribution) axes[0,1].hist(heights, bins50); axes[0,1].set_title(Height Distribution) axes[1,0].hist(ratios, bins50); axes[1,0].set_title(Aspect Ratio (W/H)) axes[1,1].hist(means, bins50); axes[1,1].set_title(Mean Brightness) plt.tight_layout() plt.savefig(image_stats.png, dpi150) print(f[INFO] 已保存统计图 image_stats.png) print(f[STATS] 抽样 {len(widths)} 张: avg_size{np.mean(widths):.0f}×{np.mean(heights):.0f}, favg_ratio{np.mean(ratios):.2f}, avg_brightness{np.mean(means):.1f}) if __name__ __main__: IMG_DIR 行人数据集2000张-yolov5格式数据集/images analyze_images(IMG_DIR)关键指标解读avg_size 640×480说明大量图像过小YOLOv5 默认输入为 640小图会被拉伸失真建议统一 resize 到 1280×720 再裁剪avg_ratio若集中在 1.7~1.8手机竖拍或 0.5~0.6监控横拍说明场景单一泛化性差avg_brightness 80整体偏暗夜间场景需额外加 brightness augment 200则过曝易丢失细节。3.2 分析目标尺度分布识别小目标瓶颈YOLOv5 对小目标检测能力弱若数据集中 60% 的行人 bounding box 宽度 32像素在 640 输入下仅占 5%则必须启用mosaicmulti-scale训练并在train.py中显式开启--rect矩形推理提升小目标 recall。# analyze_bbox_scale.py import numpy as np from pathlib import Path def analyze_bbox_sizes(label_dir, img_dir): sizes [] # 存储所有 bbox 的 width_px for lbl_path in Path(label_dir).glob(*.txt): img_path Path(img_dir) / f{lbl_path.stem}.jpg if not img_path.exists(): img_path Path(img_dir) / f{lbl_path.stem}.jpeg if not img_path.exists(): continue # 读取图像尺寸 from PIL import Image try: w_img, h_img Image.open(img_path).size except: continue # 解析 label with open(lbl_path, r) as f: for line in f: parts line.strip().split() if len(parts) ! 5: continue try: x, y, w_norm, h_norm map(float, parts[1:]) w_px int(w_norm * w_img) if w_px 0: sizes.append(w_px) except: continue if not sizes: print([WARN] 未解析到任何有效 bbox) return sizes np.array(sizes) print(f[BBOX STATS] 总共 {len(sizes)} 个行人框) print(f • 最小宽度: {sizes.min()}px, 最大宽度: {sizes.max()}px) print(f • 32px 占比: {np.sum(sizes32)/len(sizes)*100:.1f}%) print(f • 16px 占比: {np.sum(sizes16)/len(sizes)*100:.1f}%) # 绘制宽度分布 plt.figure(figsize(8,4)) plt.hist(sizes, bins100, range(0,200)) plt.axvline(32, colorr, linestyle--, label32px threshold) plt.xlabel(Bounding Box Width (pixels)) plt.ylabel(Count) plt.legend() plt.title(Pedestrian BBox Width Distribution) plt.savefig(bbox_width_dist.png, dpi150) if __name__ __main__: LABEL_DIR 行人数据集2000张-yolov5格式数据集/labels IMG_DIR 行人数据集2000张-yolov5格式数据集/images analyze_bbox_sizes(LABEL_DIR, IMG_DIR)血泪经验当32px占比 40%单纯调--img 1280不够必须配合--hyp hyp.scratch-low.yaml降低学习率--cache缓存增强后图像--evolve超参进化否则小目标 recall 会崩到 0.2 以下。4. 修复与增强用 3 个脚本把“残缺数据集”变成可训样本确认数据集存在坐标越界、小目标密集、图像尺寸不一等问题后不能删光重来——2000 张已是有限资源。我们用三个精准脚本做最小干预修复。4.1 修复越界坐标clip warn不丢样本YOLOv5 训练时若遇到x_center1会静默跳过该样本导致 batch_size 实际变小、梯度不准。正确做法是clip 到 [0,1] 并记录日志供后续人工复核# fix_bboxes_clip.py import os from pathlib import Path def clip_labels(label_dir): fixed_count 0 for lbl_path in Path(label_dir).glob(*.txt): lines [] with open(lbl_path, r) as f: for line in f: parts line.strip().split() if len(parts) ! 5: lines.append(line) continue try: cls_id int(parts[0]) x, y, w, h map(float, parts[1:]) # Clip coordinates x max(0.0, min(1.0, x)) y max(0.0, min(1.0, y)) w max(0.001, min(1.0, w)) # w,h 至少 0.001避免除零 h max(0.001, min(1.0, h)) # Ensure x±w/2 and y±h/2 stay in [0,1] x max(w/2, min(1-w/2, x)) y max(h/2, min(1-h/2, y)) lines.append(f{cls_id} {x:.6f} {y:.6f} {w:.6f} {h:.6f}\n) fixed_count 1 except: lines.append(line) with open(lbl_path, w) as f: f.writelines(lines) print(f[FIXED] 共修正 {fixed_count} 行越界坐标) if __name__ __main__: LABEL_DIR 行人数据集2000张-yolov5格式数据集/labels clip_labels(LABEL_DIR)注意x max(w/2, min(1-w/2, x))这行是关键——它保证 bbox 的 left/right 边界不越界比简单clip(x,0,1)更鲁棒。执行后务必用validate_yolo_format.py再验一次。4.2 统一图像尺寸resize pad保比例不失真监控截图常为 1920×1080手机抓拍多为 1280×720直接 resize 到 640×640 会严重拉伸。正确做法是等比缩放 黑边填充保持原始长宽比# resize_with_pad.py import cv2 from pathlib import Path def resize_and_pad(img_dir, target_size640, output_dirresized_images): Path(output_dir).mkdir(exist_okTrue) for img_path in Path(img_dir).glob(*.jpg): img cv2.imread(str(img_path)) h, w img.shape[:2] scale min(target_size/w, target_size/h) new_w, new_h int(w*scale), int(h*scale) resized cv2.resize(img, (new_w, new_h)) # Pad to target_size pad_w target_size - new_w pad_h target_size - new_h padded cv2.copyMakeBorder(resized, 0, pad_h, 0, pad_w, cv2.BORDER_CONSTANT, value(0,0,0)) # 保存并更新对应 label需同步缩放坐标 out_path Path(output_dir) / img_path.name cv2.imwrite(str(out_path), padded) # 更新 label 坐标原坐标是归一化的只需按 scale 缩放后重新归一化 lbl_path Path(行人数据集2000张-yolov5格式数据集/labels) / f{img_path.stem}.txt if lbl_path.exists(): new_lbl_path Path(resized_labels) / f{img_path.stem}.txt Path(resized_labels).mkdir(exist_okTrue) with open(lbl_path, r) as f_in, open(new_lbl_path, w) as f_out: for line in f_in: parts line.strip().split() if len(parts) 5: cls_id parts[0] x, y, w, h map(float, parts[1:]) # 原图尺寸 → 新图尺寸 → 归一化到 target_size x_new x * w / target_size y_new y * h / target_size w_new w * w / target_size h_new h * h / target_size f_out.write(f{cls_id} {x_new:.6f} {y_new:.6f} {w_new:.6f} {h_new:.6f}\n) else: f_out.write(line) if __name__ __main__: IMG_DIR 行人数据集2000张-yolov5格式数据集/images resize_and_pad(IMG_DIR, target_size640, output_dirresized_images)玄学提醒target_size设为 640 是因 YOLOv5 默认输入尺寸若你用yolov5s且 GPU 显存紧张可设为 320但需同步改train.py --img 320且小目标检测效果会下降。4.3 小目标增强mosaic copy-paste不靠合成靠迁移对32px的小目标最有效增强不是random_affine而是copy-paste从大目标图中 crop 出清晰行人 patchpaste 到背景图上同时生成新 label。这比 GAN 生成更真实且无需额外模型# copy_paste_augment.py import cv2 import numpy as np import random from pathlib import Path def copy_paste_augment(img_dir, label_dir, output_dir, paste_prob0.3): Path(output_dir).mkdir(exist_okTrue) img_paths list(Path(img_dir).glob(*.jpg)) for i, src_img_path in enumerate(img_paths): if random.random() paste_prob: continue # 读取源图及 label src_img cv2.imread(str(src_img_path)) lbl_path Path(label_dir) / f{src_img_path.stem}.txt if not lbl_path.exists(): continue bboxes [] with open(lbl_path, r) as f: for line in f: parts line.strip().split() if len(parts) 5: x, y, w, h map(float, parts[1:]) # 转回像素坐标 h_img, w_img src_img.shape[:2] x1 int((x - w/2) * w_img) y1 int((y - h/2) * h_img) x2 int((x w/2) * w_img) y2 int((y h/2) * h_img) if x2-x1 20 and y2-y1 20: # 只取足够大的目标 bboxes.append((x1,y1,x2,y2)) if len(bboxes) 2: continue # 随机选一个目标 crop x1,y1,x2,y2 random.choice(bboxes) patch src_img[y1:y2, x1:x2].copy() # 随机选另一张图 paste dst_img_path random.choice(img_paths) dst_img cv2.imread(str(dst_img_path)) h_dst, w_dst dst_img.shape[:2] # 随机位置 paste避开原 bbox 区域 paste_x random.randint(0, w_dst - patch.shape[1]) paste_y random.randint(0, h_dst - patch.shape[0]) # blend roi dst_img[paste_y:paste_ypatch.shape[0], paste_x:paste_xpatch.shape[1]] blended cv2.addWeighted(roi, 0.7, patch, 0.3, 0) dst_img[paste_y:paste_ypatch.shape[0], paste_x:paste_xpatch.shape[1]] blended # 生成新 label 行 new_x (paste_x patch.shape[1]/2) / w_dst new_y (paste_y patch.shape[0]/2) / h_dst new_w patch.shape[1] / w_dst new_h patch.shape[0] / h_dst new_line f0 {new_x:.6f} {new_y:.6f} {new_w:.6f} {new_h:.6f}\n # 保存新图 新 label new_name f{dst_img_path.stem}_cp{i:03d}.jpg cv2.imwrite(str(Path(output_dir) / new_name), dst_img) with open(Path(output_dir).parent / resized_labels / f{new_name.replace(.jpg,.txt)}, a) as f: f.write(new_line) if __name__ __main__: # 注意此脚本需配合前面的 resize 步骤确保图像尺寸一致 IMG_DIR resized_images LABEL_DIR resized_labels OUTPUT_DIR augmented_images copy_paste_augment(IMG_DIR, LABEL_DIR, OUTPUT_DIR, paste_prob0.2)翻车预警paste_prob0.2意味着每 5 张图做 1 次 copy-paste避免过拟合若patch.shape[0] 32paste 后仍是小目标无效——所以前置bboxes过滤必须保留20px。5. 避坑指南YOLOv5 训行人数据集的 4 个高频翻车点注意以下问题均来自真实项目非理论假设。每个现象都附带tensorboard --logdirruns/train下可查的日志线索方便你快速定位。5.1 现象训练 100 epoch 后 val/mAP0.5 一直卡在 0.050.12loss 曲线平缓无下降原因data.yaml中nc: 1正确但names: [person]写成了[pedestrian]或[0]导致val.py加载类别名失败mAP 计算逻辑崩溃返回假阴性。解决打开runs/train/exp/weights/last.pt用torch.load(...)[model].names查看实际 names或直接grep names data.yaml确认值为[person]。5.2 现象训练中途CUDA out of memory但nvidia-smi显示显存只占 70%原因--batch-size 16时若某张图含 50 行人密集场景YOLOv5 的build_targets()函数会动态分配 anchor 匹配内存峰值显存远超静态 batch 占用。解决在train.py开头添加torch.cuda.empty_cache()或更稳妥地用--rect参数启用矩形推理强制 batch 内图像尺寸相近抑制内存尖峰。5.3 现象验证时 detect 出大量重叠框NMS 未生效conf: 0.25下仍满屏小框原因--conf 0.25是置信度过滤阈值但 NMS 的 IOU 阈值--iou 0.45若设得过高如0.7会导致同类框合并不充分若过低如0.2又会误杀。解决在val.py中临时插入print(fIOU thresh: {opt.iou})确认传入值行人检测推荐--iou 0.5比通用0.45更鲁棒。5.4 现象部署到树莓派5后FPS 仅 23CPU 占用 100%GPU 几乎闲置原因export.py导出的.pt模型未做--halfFP16或--int8量化树莓派5 的 GPUVideoCore VII对 FP32 推理支持极差。解决导出时加--half参数python export.py --weights yolov5s.pt --include onnx --halfONNX 模型再用onnxruntime的ExecutionProvider指定VulkanExecutionProvider。6. 验证与上线用 3 个命令确认模型真正可用而非“训完了就完事”训完模型只是开始。真正的“可用”意味着① 在你的硬件上跑得动② 在你的光照/角度/遮挡下检得出③ 框的位置误差 ≤ 15px对 640 输入。下面三个命令缺一不可。6.1 第一关本地快速验证 —— 用detect.py看 raw output不要直接信val.py的 mAP 数字先看模型对单张图的原始输出python detect.py \ --weights runs/train/exp/weights/best.pt \ --source 行人数据集2000张-yolov5格式数据集/images/0001.jpg \ --conf 0.3 \ --iou 0.5 \ --save-txt \ --save-conf \ --project inference_test关键动作打开inference_test/exp/0001.jpg肉眼判断框是否贴合行人轮廓查看同目录下0001.txt确认每行末尾 confidence 是否 ≥0.3--save-conf生效若框明显偏移立即停训回溯data.yaml中train/val路径是否指向修复后的resized_images而非原始images。6.2 第二关跨设备验证 —— 树莓派5 上测真实 FPS树莓派5 部署不是复制粘贴必须实测# 在树莓派5上已安装 onnxruntime 1.16 python detect_rpi.py \ --weights runs/train/exp/weights/best.onnx \ --source /dev/video0 \ --img 640 \ --conf 0.3 \ --iou 0.5 \ --device cpu \ # 先用 CPU 测 baseline --view-img # 若 CPU FPS 5则切 GPU python detect_rpi.py \ --weights runs/train/exp/weights/best.onnx \ --source /dev/video0 \ --img 640 \ --conf 0.3 \ --iou 0.5 \ --device cuda \ --view-img必调参数表参数推荐值为什么--img640树莓派5 GPU 支持最大 640×640更大则 fallback 到 CPU--conf0.35行人检测需更高置信度过滤 false positive--iou0.5平衡 precision/recall0.45在密集场景易漏检--devicecuda必须显式指定否则 onnxruntime 默认用 CPU6.3 第三关业务层验证 —— 用precision-recall curve定义“可用阈值”mAP 是平均值但业务需要确定性比如“要求 95% 召回率下precision ≥ 80%”。这就需要 PR 曲线# plot_pr_curve.py from utils.metrics import ap_per_class from utils.plots import plot_pr_curve import torch # 加载 val 结果需先运行 val.py --task val results torch.load(runs/val/exp/results.pkl) # 由 val.py 自动生成 p, r, ap, f1, ap_class ap_per_class(results[stats]) plot_pr_curve(ap, p, r, Path(pr_curve.png))操作指引运行python val.py --weights runs/train/exp/weights/best.pt --data data.yaml --task val --save-json脚本会生成results.pkl运行上述代码得pr_curve.png在图中找到Recall0.95对应的Precision值若0.8说明当前模型达不到业务要求需回退到第 4 步增强小目标或增加--hyp hyp.finetune.yaml微调。我带过的 7 个项目里有 4 个卡在这一关——不是模型不行而是业务方口头说的“95%召回”没量化成 PR 曲线上的点。后来我们养成习惯每次训完模型第一件事就是画 PR 曲线把“可用”二字钉死在坐标轴上。希望帮到你。本文还有配套的精品资源点击获取