Open AccessPosted Content
Contrastive Learning for Sequential Recommendation
TL;DR: Zhang et al. as mentioned in this paper proposed a multi-task model called CL4SRec, which not only takes advantage of the traditional next item prediction task but also utilizes the contrastive learning framework to derive self-supervision signals from the original user behavior sequences.
read more
Abstract: Sequential recommendation methods play a crucial role in modern recommender systems because of their ability to capture a user's dynamic interest from her/his historical interactions. Despite their success, we argue that these approaches usually rely on the sequential prediction task to optimize the huge amounts of parameters. They usually suffer from the data sparsity problem, which makes it difficult for them to learn high-quality user representations. To tackle that, inspired by recent advances of contrastive learning techniques in the computer version, we propose a novel multi-task model called \textbf{C}ontrastive \textbf{L}earning for \textbf{S}equential \textbf{Rec}ommendation~(\textbf{CL4SRec}). CL4SRec not only takes advantage of the traditional next item prediction task but also utilizes the contrastive learning framework to derive self-supervision signals from the original user behavior sequences. Therefore, it can extract more meaningful user patterns and further encode the user representation effectively. In addition, we propose three data augmentation approaches to construct self-supervision signals. Extensive experiments on four public datasets demonstrate that CL4SRec achieves state-of-the-art performance over existing baselines by inferring better user representations.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Multi-level sequence denoising with cross-signal contrastive learning for sequential recommendation
Xiaofei Zhu,Liang Li,Weidong Liu,Xin Luo +3 more
1
FedOCD: A One-Shot Federated Framework for Heterogeneous Cross-Domain Recommendation
Liu Zhi,HE Xiao-hua,Xulin Ma,Li Wang,Guojiang Shen,Xiangjie Kong +5 more
1
•Posted Content
Memory Augmented Multi-Instance Contrastive Predictive Coding for Sequential Recommendation
Ruihong Qiu,Zi Huang,Hongzhi Yin +2 more
TL;DR: In this article, a memory augmented multi-instance contrastive predictive coding scheme is proposed to improve the long-term preference capture in sequential recommender models, where the memory module is designed to augment the auto-regressive prediction to enable a flexible and general representation of the encoded preference.
1
Enhancing Recommendation with Denoising Auxiliary Task
P. L. Liu,Linan Zheng,Jiale Chen,Guangfa Zhang,Jing Wang,Jinyun Fang +5 more
- 25 Sep 2024
TL;DR: This study proposes ATJT, a self-supervised method to enhance recommender systems by reweighting noisy user interaction sequences, improving model performance by 10-15% on three datasets through joint training with a noise recognition model.
A Novel Efficient Unclick Behavior Modeling Framework for Click-Through Rate Prediction
01 Oct 2022
TL;DR: Lu et al. as mentioned in this paper proposed an efficient unclick behavior modeling framework (UBM) to model the implicit negative feedback based on the click behavior modeling to learn users' complete and unbiased preferences for CTR prediction.
1
References
•Proceedings Article
Adam: A Method for Stochastic Optimization
Diederik P. Kingma,Jimmy Ba +1 more
- 01 Jan 2015
TL;DR: This work introduces Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments, and provides a regret bound on the convergence rate that is comparable to the best known results under the online convex optimization framework.
138.5K
Long short-term memory
TL;DR: A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
99K
•Posted Content
Adam: A Method for Stochastic Optimization
Diederik P. Kingma,Jimmy Ba +1 more
TL;DR: In this article, the adaptive estimates of lower-order moments are used for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimate of lowerorder moments.
82.5K
•Posted Content
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
TL;DR: A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
81.7K
Attention Is All You Need
Ashish Vaswani,Noam Shazeer,Niki Parmar,Jakob Uszkoreit,Llion Jones,Aidan N. Gomez,Łukasz Kaiser,Illia Polosukhin +7 more
- 01 Jan 2017
Abstract: The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
51.8K