Random forest classification for volcanogenic massive sulfide mineralization in the Rouyn-Noranda Area, Quebec.
In this paper we use GIS processing techniques and the random forest (RF) algorithm to assess the mineral potential of the Noranda district.
Random forest (RF) classification was applied to 37 predictor maps (vectors to mineralization) producing a Mineral Prospectivity Map (MPM) for volcanogenic massive sulfide (VMS) mineralization in the Noranda District, Abitibi subprovince, which is host to ∼20 VMS deposits and numerous subeconomic occurrences. The predictor maps were created using geological, geochemical, and geophysical data, and the known VMS deposits were used to train the RF classifier. The RF model was applied on two regions of interest (ROI) to investigate the effect of different sized ROIs on the results. Five sets of balanced and unbalanced training data were used in the experiments to examine the sensitivity of RF to different sets of training data. The probability/prospectivity maps showed very high success rate of classification with regard to training data particularly for the training sets with higher ratio of non-deposits/deposits. Accuracy assessment of classification maps using cross-validation also showed very high accuracy when assessed by the training data in all experiments again giving a higher accuracy for the training sets with higher ratios of non-deposits/deposits, however the later showed very low accuracy when applying k-fold cross-validation or when assessed by 15 VMS showings used as the test data indicating an overfitting problem with the unbalanced data. In most of the experiments, principal component 4 (PC4) of geochemical analysis, proximity to synvolcanic tonalite-trondhjemite-granodiorite (TTG) intrusions, lithology, and sericite alteration maps showed higher predictive power. The importance of individual variables changed when the ROI was changed to a smaller area suggesting that RF is sensitive to the study area.
Datasets available for download
-
Document LinkHTML
Linked dataset resource
Additional Info
| Field | Value |
|---|---|
| Last Updated | March 25, 2026, 02:05 (UTC) |
| Created | March 25, 2026, 02:05 (UTC) |
|
Domain / Topic
Domain or topic of the dataset being cataloged.
|
|
|
Title
Title for the Dataset.
|
Random forest classification for volcanogenic massive sulfide mineralization in the Rouyn-Noranda Area, Quebec. |
|
Description
A description of the dataset.
|
In this paper we use GIS processing techniques and the random forest (RF) algorithm to assess the mineral potential of the Noranda district. Random forest (RF) classification was applied to 37 predictor maps (vectors to mineralization) producing a Mineral Prospectivity Map (MPM) for volcanogenic massive sulfide (VMS) mineralization in the Noranda District, Abitibi subprovince, which is host to ∼20 VMS deposits and numerous subeconomic occurrences. The predictor maps were created using geological, geochemical, and geophysical data, and the known VMS deposits were used to train the RF classifier. The RF model was applied on two regions of interest (ROI) to investigate the effect of different sized ROIs on the results. Five sets of balanced and unbalanced training data were used in the experiments to examine the sensitivity of RF to different sets of training data. The probability/prospectivity maps showed very high success rate of classification with regard to training data particularly for the training sets with higher ratio of non-deposits/deposits. Accuracy assessment of classification maps using cross-validation also showed very high accuracy when assessed by the training data in all experiments again giving a higher accuracy for the training sets with higher ratios of non-deposits/deposits, however the later showed very low accuracy when applying k-fold cross-validation or when assessed by 15 VMS showings used as the test data indicating an overfitting problem with the unbalanced data. In most of the experiments, principal component 4 (PC4) of geochemical analysis, proximity to synvolcanic tonalite-trondhjemite-granodiorite (TTG) intrusions, lithology, and sericite alteration maps showed higher predictive power. The importance of individual variables changed when the ROI was changed to a smaller area suggesting that RF is sensitive to the study area. |
|
Tags / Keywords
Keywords/tags categorizing the dataset.
|
|
|
Format (CSV, XLS, TXT, PDF, etc)
File format of the dataset.
|
|
|
Dataset Size
Dataset size in megabytes.
|
0.12 |
|
Metadata Identifier
Metadata identifier – can be used as the unique identifier for catalogue entry
|
|
|
Published Date
Published date of the dataset.
|
2023-09-21 |
|
Time Period Data Span (start date)
Start date of the data in the dataset.
|
|
|
Time Period Data Span (end date)
End date of time data in the dataset.
|
|
|
GeoSpatial Area Data Span
A spatial region or named place the dataset covers.
|
| Field | Value |
|---|---|
|
Access category
Type of access granted for the dataset (open, closed, service, etc).
|
public |
|
License
License used to access the dataset.
|
Creative Commons Attribution |
|
Limits on use
Limits on use of data.
|
|
|
Location
Location of the dataset.
|
https://metalearth.geohub.laurentian.ca/datasets/75102e1e2c7146188f47ff9dda6f4be0 |
|
Data Service
Data service for accessing a dataset.
|
|
|
Owner
Owner of the dataset.
|
MetalEarth |
|
Contact Point
Who to contact regarding access?
|
|
|
Contact Point Email
The email to contact regarding access?
|
|
|
Publisher
Publisher of the dataset.
|
|
|
Publisher Email
Email of the publisher.
|
|
|
Author
Author of the dataset.
|
MetalEarth |
|
Author Email
Email of the author.
|
|
|
Accessed At
Date the data and metadata was accessed.
|
2023-09-21 |
| Field | Value |
|---|---|
|
Identifier
Unique identifier for the dataset.
|
75102e1e2c7146188f47ff9dda6f4be0 |
|
Language
Language(s) of the dataset
|
English |
|
Link to dataset description
A URL to an external document describing the dataset.
|
https://metalearth.geohub.laurentian.ca/datasets/75102e1e2c7146188f47ff9dda6f4be0 |
|
Persistent Identifier
Data is identified by a persistent identifier.
|
|
|
Globally Unique Identifier
Data is identified by a persistent and globally unique identifier.
|
|
|
Contains data about individuals
Does the data hold data about individuals?
|
|
|
Contains data about identifiable individuals
Does the data hold identifiable data about individual?
|
|
|
Contains Indigenous Data
Does the data hold data about Indigenous communities?
|
|
|
Portal Type
Platform type of the source portal.
|
| Field | Value |
|---|---|
|
Version
Version of the datatset
|
None |
|
Source
Source of the dataset.
|
None |
|
Version notes
Version notes about the dataset.
|
|
|
Is version of another dataset
Link to dataset that it is a version of.
|
|
|
Other versions
Link to datasets that are versions of it.
|
|
|
Provenance Text
Provenance Text of the data.
|
|
|
Provenance URL
Provenance URL of the data.
|
|
|
Temporal resolution
Describes how granular the date/time data in the dataset is.
|
|
|
GeoSpatial resolution in meters
Describes how granular (in meters) geospatial data is in the dataset.
|
|
|
GeoSpatial resolution (in regions)
Describes how granular (in regions) geospatial data is in the dataset.
|
| Field | Value |
|---|---|
|
Indigenous Community Permission
Who holds the Indigenous Community Permission. Who to contact regarding access to a dataset that has data about Indigenous communities.
|
|
|
Community Permission
Community permission (who gave permission).
|
|
|
The Indigenous communities the dataset is about
Indigenous communities from which data is derived.
|
| Field | Value |
|---|---|
|
Number of data rows
If tabular dataset, total number of rows.
|
|
|
Number of data columns
If tabular dataset, total number of unique columns.
|
|
|
Number of data cells
If tabular dataset, total number of cells with data.
|
|
|
Number of data relations
If RDF dataset, total number of triples.
|
|
|
Number of entities
If RDF dataset, total number of entities.
|
|
|
Number of data properties
If RDF dataset, total number of unique properties used by the triples.
|
|
|
Data quality
Describes the quality of the data in the dataset.
|
|
|
Metric for data quality
A metric used to measure the quality of the data, such as missing values or invalid formats.
|
0 Comments