Akane: Perplexity-Guided Time Series Data Cleaning
Summary: Akane reframes time-series cleaning as perplexity minimization: exploit recurrent patterns like token n-grams, then pick edits under a cleaning budget to lower sequence perplexity. Key novelty is perplexity-guided dirty-point detection/repair with a 4-phase framework plus budget selection and pattern-aggregation heuristics. (summarized by gpt-5.4-mini on May 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Xiaoyu Han (Fudan University)
- 2. Haoran Xiong (Fudan University)
- 3. Zhenying He (Fudan University)
- 4. Peng Wang (Fudan University)
- 5. Chen Wang (Tsinghua University)
- 6. X. Sean Wang (Fudan University)
BibTeX Citation
@inproceedings{han_sigmod24,
title = {{Akane: Perplexity-Guided Time Series Data Cleaning}},
author = {Han, Xiaoyu and Xiong, Haoran and He, Zhenying and Wang, Peng and Wang, Chen and Wang, X. Sean},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3654993},
url = {https://dl.acm.org/doi/10.1145/3654993},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,372 | From Suspicious Errors to Valid Data: On Repairing Spatio-Temporal Data via Spatial and Temporal Dependencies | 2026 | SIGMOD | 5.093636e-05 |
| 10,500 | SHoTClean: Bridging Soft and Hard Constraints for Multivariate Time Series Cleaning | 2026 | SIGMOD | 5.093636e-05 |
| 10,784 | The Best of Both Worlds: On Repairing Timestamps and Attribute Values for Multivariate Time Series | 2025 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,966 | UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning Workflow | 2025 | VLDB |
| 2 | 7,880 | Learning Over Dirty Data Without Cleaning | 2020 | SIGMOD |
| 3 | 2,995 | Time Series Data Cleaning: From Anomaly Detection to Anomaly Repairing | 2017 | VLDB |
| 4 | 9,697 | Clean4TSDB: A Data Cleaning Tool for Time Series Databases | 2024 | VLDB |
| 5 | 582 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB |
| 6 | 10,500 | SHoTClean: Bridging Soft and Hard Constraints for Multivariate Time Series Cleaning | 2026 | SIGMOD |
| 7 | 5,794 | ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning | 2016 | SIGMOD |
| 8 | 9,699 | MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series Data | 2024 | VLDB |
| 9 | 5,175 | Cleanits: A Data Cleaning System for Industrial Time Series | 2019 | VLDB |
| 10 | 10,353 | Cleaning Time Series under Seasonal and Trend Constraints | 2026 | SIGMOD |