Deep learning will certainly play a crucial role in automating and improving insurers’ day-to-day operations that involve image, text and audio processing. In the medium term, however, it is not expected to have a deep impact on insurers’ core business of risk management.
This short paper shares our point of view on the use of deep learning technologies in the insurance sector.
What deep learning is, and a little history
Deep learning is a family of machine learning techniques inspired by neural networks, involving multiple layers of models. The term is something of a buzzword (a rebranding of neural networks, now with more layers) and covers a wide range of algorithms, among them convolutional neural networks, recurrent neural networks and long short-term memory networks.
Its roots go back to the mid-20th century. The first algorithmic neural network, inspired by biology, was Rosenblatt’s perceptron in 1957.
The first big step came in the 1980s, when a meaningful framework for optimising parameters, gradient methods, was finally applied to neural networks. At that stage, the efficiency of these new methods was far from proven.
It took until the 2010s for a double revolution: in the volume of data (with the release of the ImageNet dataset) and in computing power (Nvidia CUDA GPUs reaching a trillion operations per second). Deep learning then started to beat every other method in image processing, and progressively in natural language processing too.
A revolution followed, spreading at incredible speed, in roughly a year and a half.
How does it work?
To make a long story short, understanding deep learning comes down to two aspects: architecture and back-propagation.
The architecture of a neural network involves many neurons interacting with each other in successive layers. These represent as many parameters (number of layers, types of interaction, weights of each cell) that the machine will learn by confronting this architecture with real-world data.
Then comes back-propagation. The idea is very simple: every time a new input enters the network, the output produces an error (measured by a loss function). This error is propagated back so that each neuron is assigned a contribution to it. These contributions are used to optimise the weights by gradient descent, ultimately minimising the loss function.
The real magic is that in deep learning, everything is optimised simultaneously.
These are the high-level principles; the purpose here is not to dig into the technical details, but to focus on the key elements that drive success in deep learning.
In practice, what does it take to make it work?
- A large volume of data: with hundreds of thousands of parameters to calibrate, the number of observations must be far larger.
- The right kind of input data, where the added value is proven. This is typically unstructured data: images, text and audio. The intuition is that complex, deep neural networks add the most value when the underlying variables are hard to infer, which is typically the case for images, text and audio.
- Deep learning has shown fewer significant breakthroughs over standard machine learning algorithms in other applications with structured data.
Two standard examples illustrate this: images and text.
One of the major families of deep learning architectures for image processing is convolutional neural networks. They identify and focus on different parts of an image, then assemble them to classify it. These methods learn contours, shapes, positions, contrast levels and so on. This is the high-value part; the final decision then relies on relatively simple rules.
Whenever data has a historical or temporal dimension, the architectures used are recurrent neural networks, among which one of the most popular is the LSTM (long short-term memory).
They contain loops that let them take the past into account, in the short and long term: information becomes persistent. The most famous example is predicting a word in a sentence. A memory effect is needed to infer the last word of the sentence “I have been living in Germany for five years… I am fluent in German”. A standard model could learn that the word after “fluent in” is the name of a language; to predict that it is German, the model needs memory, because “Germany” appears earlier in the sentence.
LSTMs can also be seen as the deep learning equivalent of econometric time series models. We strongly believe there are still large innovation opportunities in these approaches, both in finance, where time series are everywhere, and in insurance, where time series and space-time series are becoming omnipresent (more on this below).
An El Dorado for insurance?
Deep learning has given, and is still giving, a new boost to artificial intelligence. But it will not solve every problem insurers face. A closer look shows where deep learning can be expected to add little or much value in insurance.
To date, no one can fully explain why deep learning performs so well on so many use cases, so it does not natively lend itself to interpretation.
Yet one of an insurer’s main tasks is to understand risk and its drivers, in order to prevent and manage it as well as possible.
Interpretability is required from several perspectives: regulation, audit, decision-making and stability. It is an active field of investigation in insurance, in large groups as well as insurtech start-ups. The idea is to take the outcome of deep learning and explain it with an interpretable model… but this is only the beginning of the story.
This briefly explains why deep learning has no straightforward use in the pure risk applications insurers face.
There are, however, many other applications in insurance, and this is a very active field. Why?
Insurance is an information business, built on flows of information and data, and a large part of that data today consists of images, text and audio. As mentioned above, the added value of deep learning is clear for these types of unstructured data. The main areas where deep learning is affecting the insurance business are probably:
In the short term:
- claims, e.g. assessing claim costs from images;
- fraud, detected in images, speech and text;
- customer experience and new services: satisfaction, instant quotes;
- augmented intelligence in robotic process automation (which may shift operational risk);
- parametric insurance for agriculture, farming and any weather-sensitive activity.
And even more in the longer term:
- all the data generated by the IoT, for example space-time series for connected health or telematics.
These applications are numerous. They are designed to improve service, and therefore customer satisfaction, as well as efficiency.
In conclusion, deep learning will clearly not meet all of insurers’ data analytics needs in the short term, since their job is mainly to understand risks and interpret the related models and algorithms. Today, deep learning is not well understood, and therefore does not yet produce the models they need.
However, deep learning will be used massively in insurance, in the short and long term, wherever there are images, text, audio and IoT data.
This raises a major underlying debate about the future and role of human employees. What will be the real impact on insurance jobs? Will employees be massively replaced, or will the roles humans play shift?
