Systems that can create a summary from conversations are much in demand in customer service, news rooms, and for virtual assistants. Research in the International Journal of Applied Pattern Recognition could improve processing speed and cut computational costs by using a new compression technique that avoids the lengthy retraining normally required after algorithm optimisation.
The researchers propose a framework that cuts the summarisation model’s computational workload by 30 to 40 per cent. It does so without compromising accuracy and works about one and a half times as fast as current methods. The team tested their system against five benchmark datasets and sustained performance to within 1 per cent of standard evaluation measures, including BLEU, ROUGE, METEOR and PARENT. Those evaluation tools compare computer-generated summaries against human-written references.
The team explains that they use structured pruning instead of retraining the entire model. This allowed them to optimise the model to meet a computational constraint known as a FLOPs budget, where FLOPs (floating-point operations) represent the amount of computation needed for the processing.
The optimisation process takes less than six minutes, compared with 15 to 93 hours for conventional retraining-based methods. However, the researchers note a sharp decline in performance if the FLOPs budget is too tight, specifically below 50 per cent of the original model. This, they explain, suggests limits to current compression techniques.
Liu, Y., Zhang, Y., Sun, H. and Zhang, X. (2026) ‘Efficient post-training pruning of dialogue summarisation based on sensitivity evaluation’, Int. J. Applied Pattern Recognition, Vol. 8, No. 3, pp.1–18.
News media may use this press release as source material, in whole or in part, provided the content is not materially misrepresented. A link back to the original article is appreciated.