This repository is the cregit-codev fork of the original cregit
project.
The original upstream repository is available at https://github.com/cregit/cregit.
Requires Nix and
devenv; everything else (JDKs, sbt,
srcml, ctags, cargo, the Perl modules) is pinned by devenv.nix.
git clone https://github.com/ccsl-codev/cregit.git
cd cregit
devenv shell # enter the pinned toolchain
./run_pipeline_process.sh --repo-url https://github.com/OWNER/REPO.gitThe script builds anything missing first, then runs the whole pipeline on the
repository you point it at — e.g. --repo-url https://github.com/jqlang/jq.git
makes a nice small C demo. No target repo in mind? Validate the install by
running cregit on itself — about two minutes end to end (plus the one-time
build on the first run):
./run_pipeline_process.sh --repo-url https://github.com/ccsl-codev/cregit.git --mask '\.(c|cpp|hpp|java|rs)$'Tip: with direnv installed, run direnv allow once in
the checkout — the repo ships an .envrc, so the devenv shell then
activates automatically whenever you cd in, and typing devenv shell is no
longer needed.
By default every file the tokenizer can parse is tokenized — C, C++, Java and
Rust, case-insensitively. The default mask is derived from the tokenizer's own
extension table, so it cannot drift from it; print it with perl tokenize/fileMask.pl. Pass --mask to narrow it to one language. .am/.ac
are routed to the m4 parser but are not in the default mask, because that parser
mis-lexes real autotools quoting (see %MASKED_LANGUAGES in
tokenize/CregitLanguages.pm). The browsable per-file
HTML views land in the sibling directory ../cregit-files/html. See
How to use for all flags and outputs.
- Code is written in Scala, C++, Rust and Perl.
- Platform: Linux x86_64 or macOS arm64 —
devenvbuilds thesrcmlparser from a pinned develop revision (nix/srcml.nix).
| srcml | https://www.srcml.org/ | Make sure srcml is in path |
| ctags | https://github.com/universal-ctags | Make sure ctags is in path |
The tokenization step is now provided by the blobExec sbt module
(consumes upstream com.madgag:bfg-library from Maven Central). It replaces the
previous dmgerman/bfg-repo-cleaner@blobexec fork. See blobExec/README.md.
For each module, its dependencies are documented in their corresponding README file.
As an example, on Debian 9 the following packages must be installed:
cmake libarchive-dev libxml++2.6-dev libxml2-dev libcurl4-openssl-dev libxslt1-dev libboost-all-dev libantlr-dev libssl-dev libxerces-c-dev exuberant-ctags libdbi-perl libjgit-java libhtml-fromtext-perl libset-scalar-perl libdbd-sqlite3-perl
The pipeline script builds any missing artifact automatically before a run.
To build everything explicitly (inside devenv shell):
./run_pipeline_process.sh --build-onlyThis builds, in dependency order:
| artifact | module | toolchain |
|---|---|---|
tokenize/srcMLtoken/srcml2token |
srcML tokenizer | gcc + xerces-c + libsrcml |
tokenize/rustTokenizer binary |
Rust tokenizer | cargo |
blobExec/target/scala-2.13/blobExec-0.1.0-assembly.jar |
tokenization driver | sbt, JDK 21 |
{slickGitLog,persons,remapCommits}/target/scala-2.10/*-one-jar.jar |
history / persons / remap tools | sbt 0.13, JDK 8 |
To build a single module manually: sbt assembly in blobExec;
sbt --java-home "$LEGACY_JAVA_HOME" one-jar in slickGitLog, persons or
remapCommits (they are Scala 2.10 and do not build on a modern JDK); make
in tokenize/srcMLtoken and tokenize/rustTokenizer. Build srcml2token
before running the pipeline or the Perl test suite — the tokenizer shells out
to it.
Run the test suites inside the pinned development environment (devenv shell).
The commands below mirror the required checks in GitHub Actions:
cd blobExec && sbt -batch test assembly
cd ../tokenize/srcMLtoken && make && make test
cd ../rustTokenizer && make && make test
cd ../..
bash tests/test_token_worker.sh
prove tests/t
pytest -q
for t in tokenize_gate pipeline_workdir_guard ensure_artifacts retokenize_passthrough reblame_passthrough; do
bash "tests/$t.sh"
done
for module in slickGitLog persons remapCommits; do
(cd "$module" && sbt --java-home "$LEGACY_JAVA_HOME" -batch test one-jar)
doneThe Perl tests create temporary Git repositories and SQLite databases. The
tokenizeSrcMl tests require srcml2token, so build the C++ tokenizer before
running prove.
run_pipeline_process.sh is the driver for the whole pipeline: it clones the
target repository, tokenizes it (rewriting each matched blob to its
token-level representation), builds the history and persons databases, blames
every tokenized file, generates the HTML views and writes a unified Parquet
dataset (see generate_dataset/DATASET.md).
# smoke test — cregit on itself (validates the install, ~2 min):
./run_pipeline_process.sh --repo-url https://github.com/ccsl-codev/cregit.git --mask '\.(c|cpp|hpp|java|rs)$'
# a small C project (demo-sized):
./run_pipeline_process.sh --repo-url https://github.com/jqlang/jq.git
# a Java project, with its own output directory:
./run_pipeline_process.sh \
--repo-url https://github.com/OWNER/REPO.git \
--mask '\.java$' \
--work ../cregit-files-REPOBlame and HTML generation process independent files concurrently. By default,
the pipeline uses up to four workers (or fewer when fewer CPUs are available).
Use --jobs N or CREGIT_JOBS=N to choose another limit.
Flags (see ./run_pipeline_process.sh --help for the full list):
| flag | meaning | default |
|---|---|---|
--repo-url |
git URL or local path of the repository to process | required |
--repo-name |
short name prefixed to the output files | derived from --repo-url |
--commit-url |
base URL for commit links in the generated HTML | <repo-url minus .git>/commit/ |
--mask |
regex of files to tokenize (C, C++, Java, Rust); quote it | perl tokenize/fileMask.pl |
--work |
working/output directory | ../cregit-files |
--memo-dir |
where to memoize tokenized blobs; outside --work it survives the full-run wipe |
<work>/memo |
--mode / --shards |
tokenizer walk mode / shard count for sharded |
pipeline / 4 |
--jobs |
concurrent blame/HTML processes | CREGIT_JOBS or up to 4 CPUs |
--reblame |
re-blame every file in step 7, replacing existing .blame output |
off — a resume skips files already blamed |
git blame runs with -C100 copy detection (blameRepo/formatBlame.pl), so a
token moved between files keeps its original author. It changes up to 25% of a
project's authors, so blame from before 2026-09-22 is not comparable with later blame.
Pass --reblame when the blame itself changed: without it step 7 keeps every
existing .blame file. Do not pass it to resume an interrupted run.
A full run starts by deleting the work directory — to keep several target
repositories side by side, give each its own --work. To resume a failed run
without starting over, pass the step number printed in the step banners (with
the same target flags), e.g. ./run_pipeline_process.sh --repo-url … 5.
That wipe also takes the memo, which is the tokenizing already done: a memo hit
returns without invoking srcml at all. Put it out of reach with --memo-dir,
one directory per repository (the memo key is a hash of the file contents and
names neither repository nor extension). The runner refuses a step-1 wipe that
would delete a memo of 10,000 entries or more; --force-clean overrides it.
4 C blobs on which srcML does not terminate (eden/fs/utils/StatTimes.h, no upstream
issue) are on a blob denylist (BlobDenylistEntries.scala). They never reach the
tokenizer, are dropped from the rewritten trees, and are counted as blobsDenylisted.
A blob whose tokenizer times out, reports a parser crash or exits non-zero is
excluded the same way without an entry: it is named on an EXCLUDED failed blob
line, tried once per run, and the run exits 0. An entry still saves the --blob-timeout (600 s by
default) that a known hang costs on each run.
Example run (cregit run on itself):
p.s.: long pauses are trimmed.
Each run records a tokenizer identity per extension; a changed one with cached rows refuses
the run (exit 3). After ./run_pipeline_process.sh --ensure-artifacts, resume at step 2 with
--retokenize rs to redo only those cached tokens; one that would change nothing exits 7.
Everything lands in the work directory (default: ../cregit-files, a sibling
of the checkout):
| path | content |
|---|---|
html/ |
per-file HTML views of token-level contributions |
<name>-dataset.parquet |
unified token/commit/author dataset (schema) |
<name>-cregit.git, <name>-cregit/ |
the tokenized ("view") repository |
<name>-original.git, <name>-original/ |
bare + working clones of the target repository |
<name>-*.db |
SQLite databases: history (original and cregit), blob map, persons |
blame/ |
per-file token blame |
pipeline.log |
full log of the run |
The pipeline script sets these itself; you only need them when invoking the
tools manually (the numbered steps inside run_pipeline_process.sh are the
reference for manual invocations):
| variable | meaning |
|---|---|
BFG_MEMO_DIR |
directory used to memoize tokenized blobs |
BFG_TOKENIZE_CMD |
tokenize command; the script routes it through tokenize/tokenize.pl, which dispatches by file extension |
CREGIT_JOBS |
concurrent blame/HTML processes when --jobs is not provided (otherwise up to four CPUs) |
LEGACY_JAVA_HOME |
JDK 8 home for the Scala 2.10 modules (provided by devenv shell) |
Contributions are welcome! Please read our contributing guide before opening an issue or pull request.
For larger changes, please open an issue first to discuss the proposal.
The license of Cregit is GPL-3.0+.
- use preferred name in html files
- customize programs to read a JSON file with configuration?
