AITopics | faster r-cnn

Collaborating Authors

faster r-cnn

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

If you are looking for an answer to the question What is Artificial Intelligence? and you only have a minute, then here's the definition the Association for the Advancement of Artificial Intelligence offers on its home page: "the scientific understanding of the mechanisms underlying thought and intelligent behavior and their embodiment in machines."

However, if you are fortunate enough to have more than a minute, then please get ready to embark upon an exciting journey exploring AI (but beware, it could last a lifetime) …

Monte Carlo Stochastic Depth for Uncertainty Estimation in Deep Learning

Müller, Adam T., Rögelein, Tobias, Stache, Nicolaj C.

arXiv.org Machine LearningApr-15-2026

The deployment of deep neural networks in safety-critical systems necessitates reliable and efficient uncertainty quantification (UQ). A practical and widespread strategy for UQ is repurposing stochastic regularizers as scalable approximate Bayesian inference methods, such as Monte Carlo Dropout (MCD) and MC-DropBlock (MCDB). However, this paradigm remains under-explored for Stochastic Depth (SD), a regularizer integral to the residual-based backbones of most modern architectures. While prior work demonstrated its empirical promise for segmentation, a formal theoretical connection to Bayesian variational inference and a benchmark on complex, multi-task problems like object detection are missing. In this paper, we first provide theoretical insights connecting Monte Carlo Stochastic Depth (MCSD) to principled approximate variational inference. We then present the first comprehensive empirical benchmark of MCSD against MCD and MCDB on state-of-the-art detectors (YOLO, RT-DETR) using the COCO and COCO-O datasets. Our results position MCSD as a robust and computationally efficient method that achieves highly competitive predictive accuracy (mAP), notably yielding slight improvements in calibration (ECE) and uncertainty ranking (AUARC) compared to MCD. We thus establish MCSD as a theoretically-grounded and empirically-validated tool for efficient Bayesian approximation in modern deep learning.

artificial intelligence, machine learning, mcsd, (17 more...)

arXiv.org Machine Learning

2604.12719

Country: Europe > Germany (0.04)

Genre: Research Report > New Finding (0.66)

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (1.00)
Information Technology > Artificial Intelligence > Representation & Reasoning > Uncertainty > Bayesian Inference (0.66)

Add feedback

R-FCN: Object Detection via Region-based Fully Convolutional Networks

jifeng dai, Yi Li, Kaiming He, Jian Sun

Neural Information Processing SystemsMar-23-2026, 09:54:46 GMT

We present region-based, fully convolutional networks for accurate and efficient object detection. In contrast to previous region-based detectors such as Fast/Faster R-CNN [7, 19] that apply a costly per-region subnetwork hundreds of times, our region-based detector is fully convolutional with almost all computation shared on the entire image. To achieve this goal, we propose position-sensitive score maps to address a dilemma between translation-invariance in image classification and translation-variance in object detection. Our method can thus naturally adopt fully convolutional image classifier backbones, such as the latest Residual Networks (ResNets) [10], for object detection. We show competitive results on the PASCAL VOC datasets (e.g., 83.6% mAP on the 2007 set) with the 101-layer ResNet. Meanwhile, our result is achieved at a test-time speed of 170ms per image, 2.5-20 faster than the Faster R-CNN counterpart.

artificial intelligence, convolutional layer, machine learning, (17 more...)

Neural Information Processing Systems

Country: Asia (0.40)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (1.00)

Add feedback

PrObeD: Proactive Object Detection Wrapper

Neural Information Processing SystemsFeb-18-2026, 00:02:00 GMT

These works are regarded as passive works for object detection as they take the input image as is. However, convergence to global minima is not guaranteed to be optimal in neural networks; therefore, we argue that the trained weights in the object detector are not optimal. To rectify this problem, we propose a wrapper based on proactive schemes, PrObeD, which enhances the performance of these object detectors by learning a signal. PrObeD consists of an encoder-decoder architecture, where the encoder network generates an image-dependent signal termed templates to encrypt the input images, and the decoder recovers this template from the encrypted images. We propose that learning the optimum template results in an object detector with an improved detection performance. The template acts as a mask to the input images to highlight semantics useful for the object detector. Finetuning the object detector with these encrypted images enhances the detection performance for both generic and camouflaged.

artificial intelligence, detector, machine learning, (18 more...)

Neural Information Processing Systems

Country:

South America > Brazil (0.04)
North America > United States > Michigan (0.04)
Asia (0.04)

Industry:

Information Technology (0.47)
Health & Medicine (0.46)

Technology:

Information Technology > Sensing and Signal Processing > Image Processing (1.00)
Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.69)

Add feedback

cb0f9020c00fc52a9f6c9dbfacc6ac58-Paper-Conference.pdf

Neural Information Processing SystemsFeb-11-2026, 22:31:30 GMT

As aresult, there is aproliferation of distinct architectures and loss functions for different vision tasks.

artificial intelligence, detection, machine learning, (16 more...)

Neural Information Processing Systems

Country: Asia > China > Jiangsu Province > Changzhou (0.04)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Machine Learning (1.00)
Information Technology > Sensing and Signal Processing > Image Processing (0.94)

Add feedback

b9009beb804fa097c04d226a8ba5102e-Supplemental.pdf

Neural Information Processing SystemsFeb-10-2026, 21:22:38 GMT

ascal voc benchmark, benchmark, parameterized ap loss, (12 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence (0.54)

Add feedback

acaa23f71f963e96c8847585e71352d6-Paper.pdf

Neural Information Processing SystemsFeb-9-2026, 19:30:53 GMT

computer vision, dataset, noun, (14 more...)

Neural Information Processing Systems

Country:

North America > United States > Hawaii > Honolulu County > Honolulu (0.04)
North America > United States > Massachusetts > Suffolk County > Boston (0.04)
North America > United States > California > Los Angeles County > Long Beach (0.04)
(2 more...)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Representation & Reasoning (1.00)
Information Technology > Artificial Intelligence > Natural Language (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.68)

Add feedback

acaa23f71f963e96c8847585e71352d6-AuthorFeedback.pdf

Neural Information Processing SystemsFeb-9-2026, 19:30:41 GMT

baseline, category, cobe, (13 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Machine Learning (0.50)

Add feedback

RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder

Neural Information Processing SystemsDec-24-2025, 08:58:16 GMT

Existing object detection frameworks are usually built on a single format of object/part representation, i.e., anchor/proposal rectangle boxes in RetinaNet and Faster R-CNN, center points in FCOS and RepPoints, and corner points in CornerNet. While these different representations usually drive the frameworks to perform well in different aspects, e.g., better classification or finer localization, it is in general difficult to combine these representations in a single framework to make good use of each strength, due to the heterogeneous or non-grid feature extraction by different representations.

bridging visual representation, object detection, representation, (9 more...)

Neural Information Processing Systems

Technology:

Information Technology > Artificial Intelligence (0.60)
Information Technology > Data Science (0.41)

Add feedback

A Physics-Constrained, Design-Driven Methodology for Defect Dataset Generation in Optical Lithography

Hu, Yuehua, Kong, Jiyeong, Shin, Dong-yeol, Kim, Jaekyun, Kang, Kyung-Tae

arXiv.org Artificial IntelligenceDec-11-2025

The efficacy of Artificial Intelligence (AI) in micro/nano manufacturing is fundamentally constrained by the scarcity of high-quality and physically grounded training data for defect inspection. Lithography defect data from semiconductor industry are rarely accessible for research use, resulting in a shortage of publicly available datasets. To address this bottleneck in lithography, this study proposes a novel methodology for generating large-scale, physically valid defect datasets with pixel-level annotations. The framework begins with the ab initio synthesis of defect layouts using controllable, physics-constrained mathematical morphology operations (erosion and dilation) applied to the original design-level layout. These synthesized layouts, together with their defect-free counterparts, are fabricated into physical samples via high-fidelity digital micromirror device (DMD)-based lithography. Optical micrographs of the synthesized defect samples and their defect-free references are then compared to create consistent defect delineation annotations. Using this methodology, we constructed a comprehensive dataset of 3,530 Optical micrographs containing 13,365 annotated defect instances including four classes: bridge, burr, pinch, and contamination. Each defect instance is annotated with a pixel-accurate segmentation mask, preserving full contour and geometry. The segmentation-based Mask R-CNN achieves AP@0.5 of 0.980, 0.965, and 0.971, compared with 0.740, 0.719, and 0.717 for Faster R-CNN on bridge, burr, and pinch classes, representing a mean AP@0.5 improvement of approximately 34%. For the contamination class, Mask R-CNN achieves an AP@0.5 roughly 42% higher than Faster R-CNN. These consistent gains demonstrate that our proposed methodology to generate defect datasets with pixel-level annotations is feasible for robust AI-based Measurement/Inspection (MI) in semiconductor fabrication.

artificial intelligence, deep learning, machine learning, (18 more...)

arXiv.org Artificial Intelligence

2512.09001

Genre: Research Report (1.00)

Industry:

Semiconductors & Electronics (1.00)
Information Technology > Hardware (0.48)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.69)

Add feedback

Filters

Collaborating Authors

faster r-cnn

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

Monte Carlo Stochastic Depth for Uncertainty Estimation in Deep Learning

R-FCN: Object Detection via Region-based Fully Convolutional Networks

PrObeD: Proactive Object Detection Wrapper

cb0f9020c00fc52a9f6c9dbfacc6ac58-Paper-Conference.pdf

b9009beb804fa097c04d226a8ba5102e-Supplemental.pdf

acaa23f71f963e96c8847585e71352d6-Paper.pdf

acaa23f71f963e96c8847585e71352d6-AuthorFeedback.pdf

13f320e7b5ead1024ac95c3b208610db-Supplemental.pdf

RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder

A Physics-Constrained, Design-Driven Methodology for Defect Dataset Generation in Optical Lithography