An improved near real-time detection YOLOv11- based model is introduced, addressing high leakage rates and localization inaccuracies in detecting dangerous driving behaviors. By integrating a multi-source dataset of 21,908 images with labels for mobile phone use, drinking, eating, and smoking, we enhance data diversity using clipping, grayscale adjustment, and blurring. The model architecture is refined by incorporating Swin Transformer to improve occlusion detection and lighting adaptation, alongside RepGFPN for optimizing multi-scale feature fusion and small target accuracy via dynamic channel allocation. Experiments show the improved Yolo11-Swin+RepGFPN model achieves an mAP@0.5 of 88.9% (3.3% improvement over baseline) and an mAP@0.5:0.95 of 50.4%, maintaining a real-time speed of 105.5 FPS. Ablation tests reveal that Swin Transformer reduces hand-motion misdetections, while RepGFPN increases recall for small targets like cigarette butts by 6.2%. This solution offers high precision and low latency for in-vehicle safety systems, balancing lightweight design with detection accuracy.