Iris provenance
The exact UCI Iris source, known variants, feature metadata, and retained attribution.
Attribution: Fisher, R. (1936). Iris [Dataset]. UCI Machine Learning Repository. doi:10.24432/C56C76.
License: CC BY 4.0, as stated on the official UCI page. The complete license text is included. The original iris.names credits R. A. Fisher as creator and Michael Marshall as donor; all original notices remain intact.
Downloaded from the official Iris archive on October 10, 2026. Extracted files preserve the archive members' bytes, including final line endings. Aner uses iris.data, not bezdekIris.data.
| Property | Value |
|---|---|
| Version | uci-iris.data-sha256-6f608b71a731 |
| Selected file | iris.data |
| File size | 4,551 bytes |
SHA-256 | 6f608b71a7317216319b4d27b4d9bc84e6abd734eda7872b71a458569e2656c0 |
| Observations / features | 150 / 4 |
| Header / delimiter | No header / comma |
| Missing or nonfinite feature fields | None in the pinned file |
| Task | Three class iris plant classification |
Feature and label order
| Zero based feature index | Aner name | Measurement | Unit |
|---|---|---|---|
| 0 | sepal_length | Sepal length | cm |
| 1 | sepal_width | Sepal width | cm |
| 2 | petal_length | Petal length | cm |
| 3 | petal_width | Petal width | cm |
The fifth CSV field is the label: Iris-setosa → 0, Iris-versicolor → 1, Iris-virginica → 2, with 50 rows per class. Original row order is retained and grouped by class.
Original and corrected records
UCI documents differences between iris.data and Fisher's published data. Aner preserves the selected variant. Record numbers below are one based; Aner source row IDs are one less.
| Record | Selected iris.data | UCI correction, present in bezdekIris.data |
|---|---|---|
| 35 | 4.9,3.1,1.5,0.1,Iris-setosa | 4.9,3.1,1.5,0.2,Iris-setosa |
| 38 | 4.9,3.1,1.5,0.1,Iris-setosa | 4.9,3.6,1.4,0.1,Iris-setosa |
Byte level record comparison confirms these are the only differing records in the two downloaded variants. The selected file contains 147 distinct complete record lines: records 10, 35, and 38 are identical, as are records 102 and 143. All are retained. Identical observations do not prove the same plant was measured: no stable specimen identifier is supplied. Row splitting does not deduplicate equal contents.
Changes: downloaded files are unchanged. Aner parses features as Float64, maps labels to the stated indices, and assigns source row IDs. It does not apply corrections, remove duplicates, normalize, or reorder observations by default.
This small historical classification example does not establish large data or medical model validity. The manifest includes hashes of the archive, metadata, and alternate file.