23 September 2026

Spark flies to organise text

A distributed computing framework can group large volumes of English text more accurately and efficiently than older methods. The system has the potential to support faster analysis across education, government, and business, according to research published in the International Journal of Intelligent Information and Database Systems, allowing blocks of related text to be automatically sorted and grouped together, or clustered.

The work addresses a practical problem in digital transformation: much of the information generated by organisations is unstructured text, whose meaning can vary with context. Conventional clustering methods are often slow or produce inconsistent groupings when confronted with large, high-dimensional datasets.

The researchers used Apache Spark, a platform for processing data across multiple computers, with an improved K-means algorithm. K-means is a method of grouping data into clusters according to similarities. The researchers modified this tool by using density peaks and maximum-minimum criteria to select better starting points for the clusters, rather than relying on random initialisation.

In tests on multiple datasets they found that the new approach improved clustering accuracy by more than 10 per cent when compared with conventional K-means approaches. The system was also stable even when they ramped up the volume of data to be processed. Indeed, they could improve parallel performance by additional computing nodes.

The framework could help organisations organise and extract information from growing bodies of English-language documents more efficiently, while providing a more consistent basis for subsequent data analysis and decision-making.

Yao, X. (2026) ‘English digital transformation algorithm for distributed big data based on Spark’, Int. J. Intelligent Information and Database Systems, Vol. 18, No. 8, pp.1–28.

News media may use this press release as source material, in whole or in part, provided the content is not materially misrepresented. A link back to the original article is appreciated.

No comments: