Browse the handbook
Modules

Iris provenance

The exact UCI Iris source, known variants, feature metadata, and retained attribution.

Attribution: Fisher, R. (1936). Iris [Dataset]. UCI Machine Learning Repository. doi:10.24432/C56C76.

License: CC BY 4.0, as stated on the official UCI page. The complete license text is included. The original iris.names credits R. A. Fisher as creator and Michael Marshall as donor; all original notices remain intact.

Downloaded from the official Iris archive on October 10, 2026. Extracted files preserve the archive members' bytes, including final line endings. Aner uses iris.data, not bezdekIris.data.

PropertyValue
Versionuci-iris.data-sha256-6f608b71a731
Selected fileiris.data
File size4,551 bytes
SHA-2566f608b71a7317216319b4d27b4d9bc84e6abd734eda7872b71a458569e2656c0
Observations / features150 / 4
Header / delimiterNo header / comma
Missing or nonfinite feature fieldsNone in the pinned file
TaskThree class iris plant classification

Feature and label order

Zero based feature indexAner nameMeasurementUnit
0sepal_lengthSepal lengthcm
1sepal_widthSepal widthcm
2petal_lengthPetal lengthcm
3petal_widthPetal widthcm

The fifth CSV field is the label: Iris-setosa → 0, Iris-versicolor → 1, Iris-virginica → 2, with 50 rows per class. Original row order is retained and grouped by class.

Original and corrected records

UCI documents differences between iris.data and Fisher's published data. Aner preserves the selected variant. Record numbers below are one based; Aner source row IDs are one less.

RecordSelected iris.dataUCI correction, present in bezdekIris.data
354.9,3.1,1.5,0.1,Iris-setosa4.9,3.1,1.5,0.2,Iris-setosa
384.9,3.1,1.5,0.1,Iris-setosa4.9,3.6,1.4,0.1,Iris-setosa

Byte level record comparison confirms these are the only differing records in the two downloaded variants. The selected file contains 147 distinct complete record lines: records 10, 35, and 38 are identical, as are records 102 and 143. All are retained. Identical observations do not prove the same plant was measured: no stable specimen identifier is supplied. Row splitting does not deduplicate equal contents.

Changes: downloaded files are unchanged. Aner parses features as Float64, maps labels to the stated indices, and assigns source row IDs. It does not apply corrections, remove duplicates, normalize, or reorder observations by default.

This small historical classification example does not establish large data or medical model validity. The manifest includes hashes of the archive, metadata, and alternate file.

Aner handbook · Guides and API reference