Empirical Risk Minimization under Random Censorship

Guillaume Ausset; Stéphan Clémençon; François Portier

Article Dans Une Revue Journal of Machine Learning Research Année : 2022

Empirical Risk Minimization under Random Censorship

, (1, 2, 3) , (4)

1
2
3
4

Guillaume Ausset

Fonction : Auteur
PersonId : 1125386

Stéphan Clémençon

Fonction : Auteur
PersonId : 174491
IdHAL : stephan-clemencon
ORCID : 0000-0002-5879-9500
IdRef : 08905203X

Département Images, Données, Signal

Signal, Statistique et Apprentissage

Institut Polytechnique de Paris

François Portier

Fonction : Auteur
PersonId : 1125387

Centre de Recherche en Économie et Statistique

Résumé

We consider the classic supervised learning problem where a continuous non-negative random label Y (e.g. a random duration) is to be predicted based upon observing a random vector X valued in R d with d ≥ 1 by means of a regression rule with minimum least square error. In various applications, ranging from industrial quality control to public health through credit risk analysis for instance, training observations can be right censored, meaning that, rather than on independent copies of (X, Y), statistical learning relies on a collection of n ≥ 1 independent realizations of the triplet (X, min{Y, C}, δ), where C is a nonnegative random variable with unknown distribution, modelling censoring and δ = I{Y ≤ C} indicates whether the duration is right censored or not. As ignoring censoring in the risk computation may clearly lead to a severe underestimation of the target duration and jeopardize prediction, we consider a plug-in estimate of the true risk based on a Kaplan-Meier estimator of the conditional survival function of the censoring C given X, referred to as Beran risk, in order to perform empirical risk minimization. It is established, under mild conditions, that the learning rate of minimizers of this biased/weighted empirical risk functional is of order O P (log(n)/n) when ignoring model bias issues inherent to plug-in estimation, as can be attained in absence of censoring. Beyond theoretical results, numerical experiments are presented in order to illustrate the relevance of the approach developed.

Mots clés

Censored data empirical risk minimization U -processes statistical learning theory survival data analysis

Domaines

Mathématiques [math] Machine Learning [stat.ML] Statistiques [math.ST] Probabilités [math.PR]

Fichier principal

19-450.pdf (569.76 Ko)

Origine : Fichiers éditeurs autorisés sur une archive ouverte

Stephan Clémençon : Connectez-vous pour contacter le contributeur

https://telecom-paris.hal.science/hal-03559365

Soumis le : dimanche 6 février 2022-16:11:30

Dernière modification le : lundi 22 avril 2024-17:20:58

Archivage à long terme le : samedi 7 mai 2022-18:09:13

Dates et versions

hal-03559365 , version 1 (06-02-2022)

Identifiants

HAL Id : hal-03559365 , version 1

Citer

Guillaume Ausset, Stéphan Clémençon, François Portier. Empirical Risk Minimization under Random Censorship. Journal of Machine Learning Research, 2022. ⟨hal-03559365⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

X INSTITUT-TELECOM GENES CNRS ENSAE CREST ENSAI X-CREST LTCI IDS S2A IP_PARIS

150 Consultations

58 Téléchargements

Empirical Risk Minimization under Random Censorship

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager