Improving Mineral Prospectivity Model Generalization: An Example from Orogenic Gold Mineralization of the Sturgeon Lake Transect, Ontario, Canada
This study tries to answer the following questions to address this problem: (i) whether using additional geologically significant labeled samples can improve MPM generalization and (ii) whether using simple binary variables instead of using multiclass and continuous variables can lessen the severity of poor generalization in MPM.
Despite the ever-increasing application of machine learning (ML) algorithms in mineral prospectivity modeling (MPM), poor generalization (over-fitting) is an issue posing impediments to ML-based MPM. This issue is partly rooted in model input variables and the paucity of mineralized zones used as labeled samples for training and validating models. This study, therefore, tries to answer the following questions to address this problem: (i) whether using additional geologically significant labeled samples can improve MPM generalization and (ii) whether using simple binary variables instead of using multiclass and continuous variables can lessen the severity of poor generalization in MPM. A dataset of orogenic gold mineralization in the Sturgeon Lake transect of Ontario, Canada, hosting 22 gold deposits and 46 gold occurrences, was exploited to define two suites of predictor variables describing orogenic gold mineralization. The original suite comprised categorical and continuous variables; however, the second set was developed by converting the first suite's variables into simple binary variables. Two experiments were conducted to answer the questions raised above; while only gold deposits were deemed labeled samples in the first experiment, the second experiment included both gold deposits and occurrences in the labeled samples. Each experiment was conducted with two sets of predictor variables, leading to four models. Comparing the bias–variance trade-off of these models enabled the authors to draw some conclusions about MPM generalization. The results of this study can provide insights into controlling the generalization of prospectivity models.
Datasets available for download
-
Document LinkHTML
Linked dataset resource
Additional Info
| Field | Value |
|---|---|
| Last Updated | March 25, 2026, 02:06 (UTC) |
| Created | March 25, 2026, 02:06 (UTC) |
|
Domain / Topic
Domain or topic of the dataset being cataloged.
|
|
|
Title
Title for the Dataset.
|
Improving Mineral Prospectivity Model Generalization: An Example from Orogenic Gold Mineralization of the Sturgeon Lake Transect, Ontario, Canada |
|
Description
A description of the dataset.
|
This study tries to answer the following questions to address this problem: (i) whether using additional geologically significant labeled samples can improve MPM generalization and (ii) whether using simple binary variables instead of using multiclass and continuous variables can lessen the severity of poor generalization in MPM.
Despite the ever-increasing application of machine learning (ML) algorithms in mineral prospectivity modeling (MPM), poor generalization (over-fitting) is an issue posing impediments to ML-based MPM. This issue is partly rooted in model input variables and the paucity of mineralized zones used as labeled samples for training and validating models. This study, therefore, tries to answer the following questions to address this problem: (i) whether using additional geologically significant labeled samples can improve MPM generalization and (ii) whether using simple binary variables instead of using multiclass and continuous variables can lessen the severity of poor generalization in MPM. A dataset of orogenic gold mineralization in the Sturgeon Lake transect of Ontario, Canada, hosting 22 gold deposits and 46 gold occurrences, was exploited to define two suites of predictor variables describing orogenic gold mineralization. The original suite comprised categorical and continuous variables; however, the second set was developed by converting the first suite's variables into simple binary variables. Two experiments were conducted to answer the questions raised above; while only gold deposits were deemed labeled samples in the first experiment, the second experiment included both gold deposits and occurrences in the labeled samples. Each experiment was conducted with two sets of predictor variables, leading to four models. Comparing the bias–variance trade-off of these models enabled the authors to draw some conclusions about MPM generalization. The results of this study can provide insights into controlling the generalization of prospectivity models.
|
|
Tags / Keywords
Keywords/tags categorizing the dataset.
|
|
|
Format (CSV, XLS, TXT, PDF, etc)
File format of the dataset.
|
|
|
Dataset Size
Dataset size in megabytes.
|
0.05 |
|
Metadata Identifier
Metadata identifier – can be used as the unique identifier for catalogue entry
|
|
|
Published Date
Published date of the dataset.
|
2023-09-21 |
|
Time Period Data Span (start date)
Start date of the data in the dataset.
|
|
|
Time Period Data Span (end date)
End date of time data in the dataset.
|
|
|
GeoSpatial Area Data Span
A spatial region or named place the dataset covers.
|
| Field | Value |
|---|---|
|
Access category
Type of access granted for the dataset (open, closed, service, etc).
|
public |
|
License
License used to access the dataset.
|
License not specified |
|
Limits on use
Limits on use of data.
|
|
|
Location
Location of the dataset.
|
https://metalearth.geohub.laurentian.ca/datasets/288ed17ba44349bf906e5bc16112dfef |
|
Data Service
Data service for accessing a dataset.
|
|
|
Owner
Owner of the dataset.
|
MetalEarth |
|
Contact Point
Who to contact regarding access?
|
|
|
Contact Point Email
The email to contact regarding access?
|
|
|
Publisher
Publisher of the dataset.
|
|
|
Publisher Email
Email of the publisher.
|
|
|
Author
Author of the dataset.
|
MetalEarth |
|
Author Email
Email of the author.
|
|
|
Accessed At
Date the data and metadata was accessed.
|
2023-09-21 |
| Field | Value |
|---|---|
|
Identifier
Unique identifier for the dataset.
|
288ed17ba44349bf906e5bc16112dfef |
|
Language
Language(s) of the dataset
|
English |
|
Link to dataset description
A URL to an external document describing the dataset.
|
https://metalearth.geohub.laurentian.ca/datasets/288ed17ba44349bf906e5bc16112dfef |
|
Persistent Identifier
Data is identified by a persistent identifier.
|
|
|
Globally Unique Identifier
Data is identified by a persistent and globally unique identifier.
|
|
|
Contains data about individuals
Does the data hold data about individuals?
|
|
|
Contains data about identifiable individuals
Does the data hold identifiable data about individual?
|
|
|
Contains Indigenous Data
Does the data hold data about Indigenous communities?
|
|
|
Portal Type
Platform type of the source portal.
|
| Field | Value |
|---|---|
|
Version
Version of the datatset
|
None |
|
Source
Source of the dataset.
|
None |
|
Version notes
Version notes about the dataset.
|
|
|
Is version of another dataset
Link to dataset that it is a version of.
|
|
|
Other versions
Link to datasets that are versions of it.
|
|
|
Provenance Text
Provenance Text of the data.
|
|
|
Provenance URL
Provenance URL of the data.
|
|
|
Temporal resolution
Describes how granular the date/time data in the dataset is.
|
|
|
GeoSpatial resolution in meters
Describes how granular (in meters) geospatial data is in the dataset.
|
|
|
GeoSpatial resolution (in regions)
Describes how granular (in regions) geospatial data is in the dataset.
|
| Field | Value |
|---|---|
|
Indigenous Community Permission
Who holds the Indigenous Community Permission. Who to contact regarding access to a dataset that has data about Indigenous communities.
|
|
|
Community Permission
Community permission (who gave permission).
|
|
|
The Indigenous communities the dataset is about
Indigenous communities from which data is derived.
|
| Field | Value |
|---|---|
|
Number of data rows
If tabular dataset, total number of rows.
|
|
|
Number of data columns
If tabular dataset, total number of unique columns.
|
|
|
Number of data cells
If tabular dataset, total number of cells with data.
|
|
|
Number of data relations
If RDF dataset, total number of triples.
|
|
|
Number of entities
If RDF dataset, total number of entities.
|
|
|
Number of data properties
If RDF dataset, total number of unique properties used by the triples.
|
|
|
Data quality
Describes the quality of the data in the dataset.
|
|
|
Metric for data quality
A metric used to measure the quality of the data, such as missing values or invalid formats.
|
0 Comments