Note · 2 min read
The data split changes what training proves
A model can work on new properties in a familiar area and fail in a different market. My CatBoost experiments test both situations.
What counts as a new property
To prepare data at a larger scale, I used Jev to label 50,000 properties in the database and build the feature dataset for CatBoost. Labelling prepares the information the model learns from; evaluation checks whether that learning transfers to other properties.
In my valuation experiments, removing a property from training was not enough to establish how the model would behave in another city. It still had examples from the same market. That answers a different question from training without any listings from that area.
I tested splits by property identity and by market. I also checked repeated listings: two posts for the same property are not independent examples. If one goes into training and the other into evaluation, the result can look better than it is.
Check what the model learned
I added features, tried several seeds and compared CatBoost with simpler calculations. In some controls, permuting the attributes while preserving missing values produced similar or better results. An improvement was insufficient to establish that the model had learned those attributes’ effects.
That changes how I interpret training. More columns can add information, noise or clues about the sample. An improvement within familiar areas does not prove generalisation to other markets.
I also separate data used to select parameters from data used to evaluate them. Once I have inspected a sample’s errors and adjusted the model in response, that sample is diagnostic; I need new data to test the improvement. Model quality depends on this separation as much as the chosen algorithm.