Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

NO@K: A Training-Free Proxy Metric for Evaluating LLM-Generated Tags in Recommender Systems

This repository contains the code for reproducing the results in:

NO@K: A Training-Free Proxy Metric for Evaluating LLM-Generated Tags in Recommender Systems Proceedings of the ACM Web Conference 2027 (WWW '27)

NO@K (Neighborhood Overlap at K) is a lightweight, training-free proxy metric that predicts downstream recommendation performance of LLM-generated tag variants without any model training. It measures whether tag-induced item neighborhoods align with collaborative filtering neighborhoods on a small stratified panel of items, requiring only minutes of CPU computation.

Repository Structure

proxytag/                     # Core Python package
├── core/                     # Data loading, model, evaluation, tag handling
├── baselines/                # 5 downstream models: AutoInt, DCNv2, KAR, LLM-Rec, UniSRec
└── analysis/                 # NO@K computation, significance tests, result aggregation
scripts/                      # Experiment scripts (see Reproducing Results below)
figures/                      # Figure generation scripts for the paper
data/                         # Datasets (not included; see Data Setup)
results/                      # Experiment results (generated by scripts)
requirements.txt
pyproject.toml
LICENSE

Setup

Requirements

  • Python >= 3.10
  • CUDA-capable GPU (for downstream model training)

Installation

pip install -e .
pip install -r requirements.txt

Data Setup

Data files are not included in this repository. Download and place them in the following structure:

data/
├── amazon-books/
│   ├── metadata.parquet          # Item metadata (columns: asin, title, description)
│   └── interactions/
│       ├── train.parquet         # Columns: userId, asin, timestamp
│       ├── val.parquet
│       └── test.parquet
├── amazon-movies/
│   ├── metadata.parquet          # Item metadata (columns: asin, title, genres)
│   └── interactions/
│       ├── train.parquet         # Columns: userId, asin, timestamp
│       ├── val.parquet
│       └── test.parquet
└── tags/
    ├── amazon-books/full/        # Generated tag parquet files (one per variant)
    └── amazon-movies/full/       # Generated tag parquet files (one per variant)

Dataset statistics:

Dataset Interactions Users Items Density
Amazon-Books 1,089,995 36,478 21,394 0.140%
Amazon-Movies 1,092,537 50,748 22,130 0.097%

Both datasets use a temporal 80/10/10 split. Amazon-Books retains items with >= 20 reviews and users with >= 10 interactions. Amazon-Movies uses a lighter filter (>= 4 reviews per item, >= 10 interactions per user).

Reproducing Results

The full pipeline proceeds in 6 steps. Each step depends on the outputs of the previous one.

Step 1: Generate Tags (32 variants per dataset)

Tags are generated via the Groq batch API. Each dataset uses 2 LLM sizes x 16 prompt variants = 32 tag variants.

The 16 prompt variants cross 4 binary axes: granularity (single/multi-word), knowledge (extract-only/use-knowledge), intent (describe/recommend), and focus (content/mood-style).

Amazon-Books (Llama 3.3 70B + Llama 3.1 8B):

python scripts/generate_tags.py \
    --input_path data/amazon-books/metadata.parquet \
    --dataset amazon-books \
    --item_id_col asin \
    --title_col title \
    --description_col description \
    --split full \
    --provider groq \
    --model llama-3.3-70b-versatile \
    --groq_mode batch

Repeat with --model llama-3.1-8b-instant for the smaller model. This generates 16 parquet files per model in data/tags/amazon-books/full/.

Amazon-Movies (GPT-OSS 120B + GPT-OSS 20B):

python scripts/generate_tags.py \
    --input_path data/amazon-movies/metadata.parquet \
    --dataset amazon-movies \
    --item_id_col asin \
    --title_col title \
    --genres_col genres \
    --split full \
    --provider groq \
    --model openai/gpt-oss-120b \
    --groq_mode batch

Repeat with --model openai/gpt-oss-20b.

Step 2: Compute NO@K Proxy Metric (Table 1, Figure 2)

Computes NO@K for all 32 tag variants per dataset and correlates with downstream NDCG@10:

python scripts/correlation_analysis.py --datasets amazon-books amazon-movies

This produces:

  • results/amazon-books_nok_vs_downstream.csv
  • results/amazon-movies_nok_vs_downstream.csv
  • Concordance and Spearman correlation tables (Table 1 in paper)

Step 3: Train Baseline Models (5 models x 2 datasets)

Trains each of the 5 downstream models without tags (no-tags baseline) and with each tag variant:

python scripts/run_all_baselines.py --datasets amazon-books amazon-movies

This trains AutoInt, DCNv2, KAR, LLM-Rec, and UniSRec. Results are saved as JSON files in results/.

Step 4: Train Hybrid Models with All Tag Variants (Table 4)

Trains the hybrid model (CF + tag embeddings) with each of the 32 tag variants:

python scripts/run_all_hybrid.py --datasets amazon-books amazon-movies

This runs 32 variants x 2 datasets = 64 training jobs. Results are saved in results/.

Step 5: Panel Size and K Sensitivity Analysis (Figure 3, Appendix)

Panel size ablation (Figure 3 — how many items are needed for a reliable proxy?):

python scripts/panel_size_ablation.py

K sensitivity (Appendix — robustness to neighborhood size):

python scripts/k_sensitivity_analysis.py

Step 6: Generate Figures

All paper figures are generated from the CSV files produced in Steps 2-5:

python figures/generate_figures.py
python figures/generate_redundancy_test.py

Output PDFs are saved to figures/.

Step 7 (Optional): Per-Bucket Tag Impact (Appendix)

Analyzes which items benefit most from tags, partitioned by popularity:

python scripts/bucket_tag_impact.py

Key Results

  • NO@K achieves 9/10 significant concordances with downstream NDCG@10 (p < 0.01), with zero inversions
  • The single non-significant case (KAR on Amazon-Books) has an extremely narrow NDCG range across variants
  • NO@K outperforms 9 alternative proxy metrics (3 CF-informed variants + 6 tag-only baselines)
  • The metric stabilizes with as few as 200 panel items and is robust across neighborhood sizes K = 5-50
  • Generating tags for the best variant costs ~$5 per dataset; NO@K evaluation takes minutes on CPU

Models

Model Type Tags Used As
AutoInt Feature interaction Additional input features
DCNv2 Feature interaction Concatenated with item features
KAR Knowledge-augmented Factual knowledge, explicitly adapted
LLM-Rec LLM-based Primary item text representation
UniSRec Sequential Item representations in sequence model

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages