Gold-Standard Financial Benchmarks

We introduce the first gold-standard financial benchmark for systematically comparing word embeddings using a financial language framework. This benchmark covers seven groups of financial analogies. Each group contains 80 analogies reaching 2660 unique analogies in total for all groups. All financial analogies are developed using the Bureau van Dijk’s Orbis database and are available for download.

In the table below, the first five groups cover publicly listed US companies, the sixth group mixes US and UK publicly listed companies, and the last group mixes US, UK, China, and Japan publicly listed companies. ‘Ticker’ is the security ticker identifier, ‘Name’ is the full name of the company, ‘City’ is the headquarters location, ‘Exchange’ is the stock exchange where the company’s share is traded, ‘Country’ is the country where headquarters is located, ‘State’ (for US companies) is the state where the headquarters is located, and finally ‘Incorporation year’ is the incorporation year of the company. To generate sufficient challenges, we chose the top 20, 10 and 5 companies from the ‘very large companies’ class for groups I-V, VI and VII, respectively. The permutation of chosen companies in each group generates 380 unique analogies for each group and 2660 analogies in total. The accuracy of each word embedding is reported for each group and all groups (overall).

The specialised FinText embeddings substantially outperform the general-purpose embeddings on finance-specific analogy tasks. The FinText specification achieves the highest accuracy in every benchmark group, with FinText Word2Vec (CBOW) attaining the highest overall top-5 accuracy of 19.59%. In comparison, the overall accuracies of Google Word2Vec and WikiNews are only 0.04% and 0.45%, respectively. Thus, the best-performing FinText specification achieves an overall accuracy approximately 490 times that of Google Word2Vec and 44 times that of WikiNews. These results indicate that domain-specific training enables FinText to capture the finance-specific relationships represented in the benchmark substantially more effectively than the general-purpose alternatives.