-
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15 -
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
Paper • 2407.17140 • Published • 2 -
D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement
Paper • 2410.13842 • Published • 6 -
DEIM: DETR with Improved Matching for Fast Convergence
Paper • 2412.04234 • Published • 2
Collections
Discover the best community collections!
Collections including paper arxiv:2304.08069
-
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15 -
Conditional DETR for Fast Training Convergence
Paper • 2108.06152 • Published -
Less is More: Focus Attention for Efficient DETR
Paper • 2307.12612 • Published • 7 -
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Paper • 2010.04159 • Published • 1
-
GLIGEN: Open-Set Grounded Text-to-Image Generation
Paper • 2301.07093 • Published • 4 -
YOLO-World: Real-Time Open-Vocabulary Object Detection
Paper • 2401.17270 • Published • 42 -
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15 -
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
Paper • 2407.17140 • Published • 2
-
Kandinsky: an Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion
Paper • 2310.03502 • Published • 78 -
Transferable and Principled Efficiency for Open-Vocabulary Segmentation
Paper • 2404.07448 • Published • 12 -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Paper • 2404.07973 • Published • 32 -
COCONut: Modernizing COCO Segmentation
Paper • 2404.08639 • Published • 30
-
Wide Residual Networks
Paper • 1605.07146 • Published • 2 -
Characterizing signal propagation to close the performance gap in unnormalized ResNets
Paper • 2101.08692 • Published • 2 -
Pareto-Optimal Quantized ResNet Is Mostly 4-bit
Paper • 2105.03536 • Published • 2 -
When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
Paper • 2106.01548 • Published • 2
-
End-to-End Object Detection with Transformers
Paper • 2005.12872 • Published • 7 -
ConsistencyDet: Robust Object Detector with Denoising Paradigm of Consistency Model
Paper • 2404.07773 • Published • 1 -
Efficient Transformer Encoders for Mask2Former-style models
Paper • 2404.15244 • Published • 1 -
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15
-
Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology
Paper • 2203.00585 • Published • 2 -
Emerging Properties in Self-Supervised Vision Transformers
Paper • 2104.14294 • Published • 4 -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
Paper • 2404.06903 • Published • 21 -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Paper • 2404.07973 • Published • 32
-
DocLLM: A layout-aware generative language model for multimodal document understanding
Paper • 2401.00908 • Published • 189 -
Visual Instruction Tuning
Paper • 2304.08485 • Published • 20 -
Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
Paper • 2403.09622 • Published • 18 -
Lumiere: A Space-Time Diffusion Model for Video Generation
Paper • 2401.12945 • Published • 86
-
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15 -
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
Paper • 2407.17140 • Published • 2 -
D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement
Paper • 2410.13842 • Published • 6 -
DEIM: DETR with Improved Matching for Fast Convergence
Paper • 2412.04234 • Published • 2
-
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15 -
Conditional DETR for Fast Training Convergence
Paper • 2108.06152 • Published -
Less is More: Focus Attention for Efficient DETR
Paper • 2307.12612 • Published • 7 -
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Paper • 2010.04159 • Published • 1
-
GLIGEN: Open-Set Grounded Text-to-Image Generation
Paper • 2301.07093 • Published • 4 -
YOLO-World: Real-Time Open-Vocabulary Object Detection
Paper • 2401.17270 • Published • 42 -
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15 -
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
Paper • 2407.17140 • Published • 2
-
End-to-End Object Detection with Transformers
Paper • 2005.12872 • Published • 7 -
ConsistencyDet: Robust Object Detector with Denoising Paradigm of Consistency Model
Paper • 2404.07773 • Published • 1 -
Efficient Transformer Encoders for Mask2Former-style models
Paper • 2404.15244 • Published • 1 -
DETRs Beat YOLOs on Real-time Object Detection
Paper • 2304.08069 • Published • 15
-
Kandinsky: an Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion
Paper • 2310.03502 • Published • 78 -
Transferable and Principled Efficiency for Open-Vocabulary Segmentation
Paper • 2404.07448 • Published • 12 -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Paper • 2404.07973 • Published • 32 -
COCONut: Modernizing COCO Segmentation
Paper • 2404.08639 • Published • 30
-
Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology
Paper • 2203.00585 • Published • 2 -
Emerging Properties in Self-Supervised Vision Transformers
Paper • 2104.14294 • Published • 4 -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
Paper • 2404.06903 • Published • 21 -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Paper • 2404.07973 • Published • 32
-
Wide Residual Networks
Paper • 1605.07146 • Published • 2 -
Characterizing signal propagation to close the performance gap in unnormalized ResNets
Paper • 2101.08692 • Published • 2 -
Pareto-Optimal Quantized ResNet Is Mostly 4-bit
Paper • 2105.03536 • Published • 2 -
When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
Paper • 2106.01548 • Published • 2
-
DocLLM: A layout-aware generative language model for multimodal document understanding
Paper • 2401.00908 • Published • 189 -
Visual Instruction Tuning
Paper • 2304.08485 • Published • 20 -
Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
Paper • 2403.09622 • Published • 18 -
Lumiere: A Space-Time Diffusion Model for Video Generation
Paper • 2401.12945 • Published • 86