Skip to content

Хранение и алгоритмы сжатия данных - ПР 2 - Spark: ORC vs Parquet - #8

Merged
timermakov merged 2 commits into
mainfrom
SaDCA-lab2-spark-orc-parquet
Oct 26, 2025
Merged

timermakov merged 2 commits into
mainfrom
SaDCA-lab2-spark-orc-parquet

Conversation

@timermakov

@timermakov timermakov commented Oct 26, 2025 •

Copy link
Copy Markdown
Collaborator

Необходимо:

  1. познакомиться с основами программирования в экосистеме Apache Spark
  2. разобраться в способах кодирования данных в форматах Parquet/ORC
  3. при помощи Apache Spark построить простой процесс преобразования сырых данных
    проекта "Фундаментальные основы систем обработки больших данных" в один из этих
    форматов (чтение - какое-либо простое произвольное преобразование - запись) и сделать
    оценку эффективности хранения - сколько места занимают и как быстро их можно читать
    В качестве справочного материала можно использовать https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html

Отчёт - SaDCA-lab2.md

@timermakov
timermakov requested review from Copilot and teyhd October 26, 2025 01:53

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR implements a Spark-based comparison of ORC and Parquet columnar storage formats for the PeaceDB project. The implementation includes a custom Spark DataSource V2 connector to read from the PeaceDatabase REST API, an ETL job to transform and write data in both formats, and tooling to analyze performance metrics.

Key changes:

  • Custom Spark DataSource V2 connector (Java) for reading documents from PeaceDB via REST API with configurable timeouts and retry logic
  • ETL job that parses JSON data into typed columns and writes to Parquet/ORC formats while measuring write/read performance
  • Python utilities for preparing news datasets and visualizing benchmark results

Reviewed Changes

Copilot reviewed 25 out of 36 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
tools/prepare_news_for_upload.py Utility to transform news dataset JSON files into PeaceDB document format
tools/plot_peacedb_columnar.py Visualization script for comparing Parquet vs ORC metrics
tools/docs/SaDCA-lab2-spark-orc-parquet/metrics.csv Benchmark results CSV data
tools/.gitignore Added ignores for NDJSON/JSONL files
spark/pom.xml Maven parent POM defining Spark 3.5.1 dependencies and build configuration
spark/peacedb-job/src/main/java/io/peacedb/job/PeacedbToColumnar.java Main Spark job comparing Parquet/ORC performance across dataset sizes
spark/peacedb-job/pom.xml Maven configuration for Spark job module
spark/peacedb-datasource/src/main/resources/META-INF/services/org.apache.spark.sql.sources.DataSourceRegister Service registration for custom DataSource
spark/peacedb-datasource/src/main/java/io/peacedb/spark/read/*.java DataSource V2 implementation for reading from PeaceDB API
spark/peacedb-datasource/src/main/java/io/peacedb/spark/model/Document.java Document model matching API schema
spark/peacedb-datasource/src/main/java/io/peacedb/spark/http/ApiClient.java HTTP client with retry logic for PeaceDB API
spark/peacedb-datasource/src/main/java/io/peacedb/spark/PeacedbTable.java Table schema definition
spark/peacedb-datasource/src/main/java/io/peacedb/spark/PeacedbDataSource.java DataSource provider implementation
spark/peacedb-datasource/pom.xml Maven configuration for datasource module
spark/README.md Documentation for building and running Spark components
spark/.gitignore Java and Spark-specific ignore patterns
SaDCA-lab2.md Lab report documenting the Spark/ORC/Parquet comparison
SaDCA-lab1.md Updated image paths for consistency
README.MD Updated controller name reference
PeaceDatabase/Core/Services/IDocumentService.cs Added Stats method interface
PeaceDatabase/Controllers/DbAndDocumentsController.cs Added stats endpoint and fixed AllDocs total count

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread PeaceDatabase/Controllers/DbAndDocumentsController.cs
@timermakov timermakov changed the title Хранение и алгоритмы сжатия данных - ПР 1 - Spark: ORC vs Parquet Хранение и алгоритмы сжатия данных - ПР 2 - Spark: ORC vs Parquet Oct 26, 2025
@timermakov
timermakov merged commit 05633ba into main Oct 26, 2025
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants