Pre-interpolation loss behavior in neural networks

Venter, Arthur Edgar William; Theunissen, Marthinus Wilhelm; Davel, Marelie Hattingh

dc.contributor.author	Venter, Arthur Edgar William
dc.contributor.author	Theunissen, Marthinus Wilhelm
dc.contributor.author	Davel, Marelie Hattingh
dc.date.accessioned	2021-03-17T15:58:53Z
dc.date.available	2021-03-17T15:58:53Z
dc.date.issued	2020
dc.identifier.isbn	978-3-030-66151-9
dc.identifier.issn	1865-0929
dc.identifier.uri	http://hdl.handle.net/10394/36914
dc.description.abstract	When training neural networks as classifiers, it is common to observe an increase in average test loss while still maintaining or improving the overall classification accuracy on the same dataset. In spite of the ubiquity of this phenomenon, it has not been well studied and is often dismissively attributed to an increase in borderline correct classifications. We present an empirical investigation that shows how this phenomenon is actually a result of the differential manner by which test samples are processed. In essence: test loss does not increase overall, but only for a small minority of samples. Large representational capacities allow losses to decrease for the vast majority of test samples at the cost of extreme increases for others. This effect seems to be mainly caused by increased parameter values relating to the correctly processed sample features. Our findings contribute to the practical understanding of a common behaviour of deep neural networks. We also discuss the implications of this work for network optimisation and generalisation.	en_US
dc.language.iso	en	en_US
dc.publisher	Springer	en_US
dc.subject	Overfitting	en_US
dc.subject	Generalization	en_US
dc.subject	Deep Learning	en_US
dc.title	Pre-interpolation loss behavior in neural networks	en_US
dc.type	Article	en_US

Files in this item

Name:: Pre-interpolation-loss-behavio ...
Size:: 1.374Mb
Format:: PDF

View/Open

This item appears in the following Collection(s)

Faculty of Engineering [1123]

Show simple item record