This is the repository for the project Keep it Private: Unsupervised Privatization of Online Text
pip install huggingface_hub
huggingface-cli download csbao/kip-dipper-large --local-dir models/dipper-large
huggingface-cli download csbao/kip-dipper-large dipper_cd_v130.bin --local-dir models./scripts/setup.shThis will create a conda environment named kip and install all dependencies.
Keep it Private performs authorship transfer by performing authorship transfer using a seq2seq model that was adversarially fine-tuned via reinforcement learning using a set of rewards (Privacy, Sense, and Soundness metrics)
The input file should be a JSONL file, one JSON object per line, with a fullText key:
{"fullText": "hi! this is the first document to privatize."}
{"fullText": "This is another text input. It can have multiple sentences..."}
$ conda activate kip
$ python src/generate.py --input_data_path ${INPUT_DATA_PATH} \
--output_path ${OUTPUT_PATH} \
--model_path models \
--model_name_to_use dipper-large \
--model_start_file ${BIN_FILE} \
--token_max_length 256 \
python src/generate.py --input_data_path {JSONFILE} \
--output_path {OUTPUT_FILE} \
--model_path models \
--model_name_to_use dipper-large \
--model_start_file models/dipper_cd_v130.bin \
--token_max_length 256input_data_path: path to the query documents to be privatizedoutput_path: file to save the privatized documentsmodel_start_file: path to trained KiP modelmodel_name_to_use: path to pre-trained base modeltoken_max_length: max cutoff length of outputrandom_seed: initialize all random seed to this valuelex_diversity: lexical diversity level (20, 40, or 60)order_diversity: order diversity level (20, 40, or 60)