SUN dataset图像数据集下载

SUN dataset数据集，有两个不错的网址：

http://vision.princeton.edu/projects/2010/SUN/ （普林斯顿大学）

http://groups.csail.mit.edu/vision/SUN/ （麻省理工学院）

普林斯顿大学的SUN数据集主页：

SUN Database: Scene Categorization Benchmark

Abstract

Scene categorization is a fundamental problem in computer vision. However, scene understanding research has been constrained by the limited scope of currently-used databases which do not capture the full variety of scene categories. Whereas standard databases for object categorization contain hundreds of different classes of objects, the largest available dataset of scene categories contains only 15 classes. In this paper we propose the extensive Scene UNderstanding (SUN) database that contains 899 categories and 130,519 images. We use 397 well-sampled categories to evaluate numerous state-of-the-art algorithms for scene recognition and establish new bounds of performance. We measure human scene classification performance on the SUN database and compare this with computational methods.

Paper

J. Xiao, J. Hays, K. Ehinger, A. Oliva, and A. Torralba.
SUN Database: Large-scale Scene Recognition from Abbey to Zoo.
IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

J. Xiao, K. A. Ehinger, J. Hays, A. Torralba, and A. Oliva.
SUN Database: Exploring a Large Collection of Scene Categories
International Journal of Computer Vision (IJCV)

Benchmark Evaluation

We use 397 well-sampled categories to evaluate numerous state-of-the-art algorithms for scene recognition and establish new bounds of performance. The results are shown in the figure on the right.

Download Figure4b in Matlab Editable Format (You can put your own curve in the figure.)

Results Visualization

We visualize the results using the combined kernel from all features for the first training and testing partition in the following webpage. For each of the 397 categories, we show the class name, the ROC curve, 5 sample traning images, 5 sample correct predictions, 5 most confident false positives (with true label), and 5 least confident false negatives (with wrong predicted label).

Recognition Results Webpage

Image Database

The database contains 397 categories SUN dataset used in the benchmark of the paper. The number of images varies across categories, but there are at least 100 images per category, and 108,754 images in total. Images are in jpg, png, or gif format. The images provided here are for research purposes only.

SUN397.tar.gz (tar.gz file, 39GB, md5sum=8ca2778205c41d23104230ba66911c7a).
Download by URLs: Dataset Image URLs

Training and Testing Partition

For the results in the paper we use a subset of the dataset that has 50 training images and 50 testing images per class, averaging over the 10 partitions in the following. To plot the curve in Figure 4(b) of the paper, we use the first n=(1, 5, 10, 20) images outof the 50 training images per class for training, and use all the same 50 testing images for testing no matter what size the training set is. (If you are using Microsoft Windows, you may need to replace / by in the following files.)

Download All Partitions (zip file).

Soucre Code for Benchmark Evaluation

Source Code Download

Scene Hierarchy

We have manually built an overcomplete three-level hierarchy for all 908 scene categories. The scene categories are arranged in a 3-level tree: with 908 leaf nodes (SUN categories) connected to 15 parent nodes at the second level (basic-level categories) that are in turn connected to 3 nodes at the first level (superordinate categories) with the root node at the top. The hierarchy is not a tree, but a Directed Acyclic Graph. Many categories such as "hayfield" are duplicated in the hierarchy because there might be confusion over whether such a category belongs in the natural or man-made sub-hierarchies.

Explore SUN Database

Kernel Matrices for SVM

Combined Training Kernel
Combined Testing Kernel
Best Single Training Kernel (HOG2x2)
Best Singe Testing Kernel (HOG2x2)
Other kernel matrices are available at THIS LINK.

Feature Matrices

The feature matrices are avialble at THIS LINK.

Human Classification Experiments

Human Confusion Matrix (Of the 13 good workers): good_workers_confusion.mat.
Overall confusion matrix and code for analysis: human_release.zip.
Mturk template for the experiment: template

DrawMe: A light-weight Javascript library for line drawing on a picture

DrawMe is a light-weight Javascript library to enable client-end line drawing on a picture in a web browser. It is targeted to provide a basis for self-define labeling tasks for computer vision researchers. It is different from LabelMe, which provides full support but fixed labeling interface. DrawMe is a Javascript library only and the users are required to write their own code to make use of this library for their specific need of labeling. DrawMe does not provide any server or server-end code for labeling, but gives the user greater flexibility for their specific need. It also comes with a simple example with Amazon Mechanical Turk interface that serializes Javascript DOM object into text for HTML form submission. The user can easily build their own labeling interface based on this MTurk example to make use for the Amazon Mechanical Turk for labeling, either using paid workers or the researchers themselves with MTurk sandbox.

Download DrawMe

——————————————————————————————我是分割线——————————————————————————————

麻省理工学院的SUN数据集主页：

Goals

The goal of the SUN database project is to provide researchers in computer vision, human perception, cognition and neuroscience, machine learning and data mining, computer graphics and robotics, with a comprehensive collection of annotated images covering a large variety of environmental scenes, places and the objects within. To build the core of the dataset, we counted all the entries that corresponded to names of scenes, places and environments (any concrete noun which could reasonably complete the phrase I am in a place, or Let’s go to the place), using WordNet English dictionary. Once we established a vocabulary for scenes, we collected images belonging to each scene category using online image search engines by quering for each scene category term, and annotate the objects in the images manually.

Scene Recognition Benchmark

To evaluate descriptors and classifiers for scene classification:

SUN397 Scene benchmark (397 scene categories), tar file (37GB, md5sum=58b3a6f1b8d6ec003458f940ada226bb) and project page for code, precomputed features, etc.

Object Detection Benchmark

The next collections contains only the fully annotated images from SUN. Each release contains the images from previous years.

SUN2012: 16,873 images, tar file (7.3GB).
Also available in PASCAL format: tar file (5.9GB). It also includes training and testing split that we recommend to follow.
See the instruction for using DPMv5 + PASCAL VOC Devkit + SUN2012.
Download the DPM v5 models trained using SUN2012

Citation

If you find this dataset useful, please cite this paper (and refer the data as SUN397, SUN2012, or SUN):

"SUN Database: Large-scale Scene Recognition from Abbey to Zoo". J. Xiao, J. Hays, K. Ehinger, A. Oliva, and A. Torralba. IEEE Conference on Computer Vision and Pattern Recognition, 2010.

To know more about the object annotation process (and the annotator), check this technical note:

"Notes on image annotation". A. Barriuso and A. Torralba. arXiv:1210.3448 [cs.CV] (unreferred).

Download Latest Dataset

You can download the raw SUN database using the LabelMe toolbox. If you do not have the latest version of the toolbox (or if you do not have the function SUNinstall.m), you should download the toolbox first:

LabelMe toolbox

To download the latest version of the database enter the Matlab commands:

>> yourpathimages = 'SUNDATABASE/Images';
>> yourpathannotations = 'SUNDATABASE/Annotations';
>> SUNinstall(yourpathimages, yourpathannotations);

The variables yourpathimages and yourpathannotations should point to the local paths where you want to download the images and annotations.

The first time that you call SUNinstall it will download the full set of images and annotations. Subsequent calls to SUNinstall will only download any new images added since the last download and the full set of annotations. If the download is interrupted the next call will not download again the images already downloaded.

If you want to download only one folder, you can specify a folder name:

>> folder = 'b/beach';
>> SUNinstall(yourpathimages, yourpathannotations, folder);

As new images are annotated everyday, you will get a slightly changing version if you download the database several times. If you are looking for a frozen copy of the database, use the links in the benchmark sections above.