Difficulty
Your best score
N/A
Gold has been considered for centuries a store of value, protecting wealth over the long term—especially during periods of inflation, currency depreciation, or economic instability. Central banks hold significant gold reserves, reflecting its importance in the global financial system.
Your role is to develop an automated system capable of predicting the gold closing price (gold close) using a complex set of financial, economic and market indicators, including stock indices, commodity prices and foreign exchange rates.
You have the following CSV file available:
Each row contains the following columns:
date – date in YYYY-MM-DD formatsp500 open, sp500 high, sp500 low, sp500 close, sp500 volume, sp500 high-low – S&P500 valuesnasdaq open, nasdaq high, nasdaq low, nasdaq close, nasdaq volume, nasdaq high-low – Nasdaq valuesus_rates_% – US reference interest rateCPI – Consumer Price Indexusd_chf, eur_usd – exchange ratesGDP – Gross Domestic Productsilver open, silver high, silver low, silver close, silver volume, silver high-low – silver valuesoil open, oil high, oil low, oil close, oil volume, oil high-low – oil valuesplatinum open, platinum high, platinum low, platinum close, platinum volume, platinum high-low – platinum valuespalladium open, palladium high, palladium low, palladium close, palladium volume, palladium high-low – palladium valuesgold open, gold high, gold low, gold close, gold volume – gold values (our target is gold close)💡 The data may have mixed granularities (daily, monthly, quarterly). Some values are missing (NaN), and the model must handle these correctly.
Build a system that can predict the gold closing price (gold close) for the rows in test.csv.
Save your predictions in a file named submission.csv with the following format:
ID,gold closeR00001,1820.35R00002,1815.92R00003,1822.10where:
ID – the unique identifier of the row from test.csvgold close – the price predicted by your systemPredictions will be compared against the real values in ground_truth.csv and evaluated using Root Mean Squared Error (RMSE):
RMSE = sqrt( sum((y_true - y_pred)^2) / n )Final scoring is based on the RMSE obtained: