Econometrics with Pre-Trained Embeddings for Unstructured Data
arXiv stat.ML1w4 min read
arXiv:2607.17378v1 Announce Type: cross Abstract: Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machine-learning models on unstructured data is costly, economists often use off-the-shelf pre-trained deep learning models developed by computer scientists to extract embeddings, which are then used as covariates in target economic analyses. Despite the popularity of this practice, its theoretical foundations remain limited. There are two main difficulties. First, the pre-trained model is usually trained on a different dataset and for a dif