

<?xml version="1.0" encoding="UTF-8"?>
<record>
  <title>A New Recognition Approach for Logical Link Blocks in Webpages</title>
  <journal>Journal of Digital Information Management</journal>
  <author>X.M. WANG, Z.D. WU, Y.N. HUANG, Q. GU</author>
  <volume>13</volume>
  <issue>2</issue>
  <year>2015</year>
  <doi></doi>
  <url></url>
  <abstract>Link block is a block structure widely
existing in webpages. Existing approaches to link blocks recognition generally suffer from two drawbacks: 1) they are designed only aiming at link blocks of physical structure, and even only aiming at specific link blocks of
block-level elements; and 2) the discovery and recognition of link blocks are based on analyzing HTML tag trees, consequently, often leading to high computing cost and
thus making them fail to deal with the diversified nonstandard webpages on the Internet. To this end, in this
paper we propose the concept of logical link blocks and then present an effective approach to discover and recognize logical link blocks from webpages. In the approach logical link blocks are recognized through scanning HTML codes and calculating the distance between adjacent links, and then two distance thresholds are used to determine the final logical link blocks. As a result, the approach not only can be free from the limits
of specific block-level link blocks, but also can greatly improve the robustness as the analysis on  HTML tag trees is no longer required. Finally, experimental results
demonstrate the effectiveness of the proposed approach, which not only provide a new way for the recognition of logical link blocks and text extraction, but also can be applied in other web information processing and mining
fields due to less demanding for particle size control of link blocks.</abstract>
</record>
