Best 10 Java Web Scraping Libraries

24 September 2026 (updated) | 13 min read

ScrapingBee is the best web scraping solution, including for Java projects, because it handles JavaScript rendering, sessions, and common anti-bot blocks for you.

This article compares the most popular Java web scraping libraries and shows when you can stick with a library versus when you should use ScrapingBee. Web scraping looks simple with an HTTP client, but real sites quickly require cookie and session handling, dynamic content support, and mitigation for CAPTCHAs, IP blocks, and rate limits.

Best Java web scraping libraries

Overview of Java Web Scraping

  1. ScrapingBee - Web scraping API with full Java support for JavaScript rendering and anti-bot bypass.
  2. jsoup - Popular library to handle any type of DOM querying and manipulation, including support for XPath and CSS selector queries.
  3. HtmlUnit - A full-fledged crawler engine with built-in support for HTML rendering and JavaScript support.
  4. Selenium - Powerful browser automation framework for Java.
  5. crawler4j - A simple crawler library for Java.
  6. Apache Nutch - Apache project for an extensible and scalable web crawler platform.
  7. Jaunt - Library with an embedded headless browser engine and native REST and JSON support.
  8. WebMagic - A framework combining HttpClient and HtmlUnit.
  9. Gecco - Lightweight crawler engine, combining jsoup and HttpClient.
  10. StormCrawler - An SDK for scalable, low-latency crawlers.

1. ScrapingBee

ScrapingBee homepage - the best web scraping API to avoid getting blocked

ScrapingBee is a comprehensive platform that's designed to make web scraping trivial. It enables users to deal with common scraping challenges, including the most demanding ones like:

  • Avoiding CAPTCHA
  • JavaScript-heavy websites
  • IP rotation
  • rate limiting
  • and more

Underneath, it uses headless browsers to mimic real user interactions.

It has dedicated support for no-code web scraping and Google search results scraping, and it can even take screenshots of the actual website rather than HTML!

Technically speaking, ScrapingBee is a platform, not a library, which makes it technology-agnostic and usable directly without any external libraries.

Here's a quick example of how to use ScrapingBee with Java:

import java.net.URI;
import java.net.URLEncoder;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.charset.StandardCharsets;

public class Main {
    public static void main(String[] args) throws Exception {

        var request = HttpRequest.newBuilder()
          .uri(URI.create("https://app.scrapingbee.com/api/v1/"
            + "?url=" + URLEncoder.encode("https://example.com", StandardCharsets.UTF_8)
            + "&extract_rules=" + URLEncoder.encode("""
            {"post-title": "h1"}
            """, StandardCharsets.UTF_8)))
          .header("Authorization", "Bearer YOUR_API_KEY")
          .GET()
          .build();

        try (var client = HttpClient.newHttpClient()) {
            var response = client.send(request, HttpResponse.BodyHandlers.ofString());
            System.out.println(response.body());
        }
    }
}

All the possible configuration options can be found in the documentation.

Ready to simplify your web scraping tasks? Sign up now to get your free API key and enjoy 1000 free credits to explore all that ScrapingBee has to offer!

2. jsoup

jsoup is one of the most popular Java libraries used for web scraping. It's lightweight and focuses mainly on raw HTML parsing and data extraction.

This section serves as a Java web scraper tutorial and Java web scraping tutorial, covering querying HTML and extracting desired data from HTML code using jsoup.

It will help you fetch webpage data, but won't handle advanced use cases like:

  • JavaScript execution
  • Advanced anti-scraping measures like IP rotation, CAPTCHA solving

Using jsoup for web scraping is straightforward. To begin, you simply need to establish a connection to the target webpage using the connect() method. Once connected, you can retrieve the HTML content and start extracting the data you need.

Here's a quick code example:

public static void main(String[] args) throws Exception {
    Document doc = Jsoup.connect("https://example.com").get();
    Elements h1Elements = doc.select("h1");
    h1Elements.forEach(h1 -> System.out.println("Title: " + h1.text()));
}

The code above uses the connect() method to retrieve a web page. The Document object represents the HTML document as a Java object, allowing you to query and extract the desired data.

However, if you just want to parse HTML, you can do that as well without making any HTTP calls:

public static void main(String[] args) {
    Document doc = Jsoup.parse("<h1>Foo</h1><h2>Bar</h2>");
    Elements h1Elements = doc.select("h1");
    h1Elements.forEach(h1 -> System.out.println("Title: " + h1.text()));
}

When scraping, using browser developer tools is essential. Open the developer tool (usually F12 or right-click > Inspect) to inspect page elements, identify HTML tags, classes, and attributes. This helps you target the correct elements for extraction.

To extract a single element, use getElementById():

Element mainHeader = doc.getElementById("main-header");
System.out.println(mainHeader.text());

This returns a single Element. Methods like getElementsByClass() or select() return multiple elements as an Elements collection, which extends ArrayList<Element>. For example, to get all the rows in a table:

Elements rows = doc.select("table tr");
for (Element row : rows) {
    System.out.println(row.text());
}

This code snippet shows how to handle multiple elements and extract all the rows from a table.

jsoup provides all the available methods for querying HTML, such as getElementById (for a singular element), getElementsByClass (for class elements), select() and selectFirst() (for CSS selectors), and traversal methods like parent(), children(), and child(). These methods help you extract the desired data from the HTML code efficiently.

Jsoup can handle malformed HTML effectively, making it suitable for static pages. Handling parsing errors is crucial, especially when dealing with poorly structured HTML.

For security, jsoup includes a Safelist-based cleaner to sanitize user-submitted content and prevent XSS attacks.

Note: Java libraries like jsoup may struggle with websites that heavily rely on JavaScript for rendering content, as they do not execute JavaScript. For such cases, consider using tools that support JavaScript execution.

You can find its sources on GitHub.

3. HtmlUnit

HtmlUnit is a "GUI-less browser for Java" with some JavaScript support. As the name suggests, it was designed for testing but will do the trick for general web scraping use cases. HtmlUnit is a GUI-less browser for Java that can execute JavaScript code and render pages like a real browser.

HtmlUnit is a good choice for scraping dynamic websites that rely heavily on JavaScript, as it can execute JavaScript code and render the page as a real browser would.

Downsides:

  • HtmlUnit does not support IP rotation, making it vulnerable to IP blocking if a website implements rate limiting.
  • It is not as fast as other libraries like jsoup, as it requires a full browser environment to execute JavaScript.

HtmlUnit can be configured to disable JavaScript and CSS rendering, which is useful for web scraping when those features are not needed. Disabling JavaScript and CSS can speed up scraping and reduce resource usage.

HtmlUnit allows you to interact with web pages by reading text, filling forms, and clicking buttons, making it suitable for automating complex interactions.

Here's a quick example of how to use HtmlUnit:

public static void main(String[] args) throws IOException {
    try (var webClient = new WebClient()) {
        HtmlPage page = webClient.getPage("https://example.com");
        HtmlAnchor anchor = page.getFirstByXPath("//a");
        if (anchor != null) {
            System.out.println("Found links:");
            System.out.printf("- '%s' -> %s%n", anchor.getVisibleText(), anchor.getHrefAttribute());
        } else {
            System.out.println("No anchors found!");
        }
    }
}
// Found links:
// - 'Learn more' -> https://iana.org/domains/example

You can find its source code on GitHub.

4. Selenium

Selenium is a browser automation framework designed for end-to-end testing but can also be leveraged for web scraping!

Controlling real browsers is Selenium's biggest advantage, but it also has a couple of downsides:

  • Scripts are often fragile and break easily when a web application's UI changes
  • Selenium uses real browsers; therefore, it's quite resource-intensive
  • It's not self-sufficient - it requires additional setup to interact with installed browsers

Here's a quick example of how to use Selenium with Java (remember about installing a browser and a driver first):

public static void main(String[] args) {
    // requires chrome and chrome-driver installed
    System.setProperty("webdriver.chrome.driver", "path/to/chromedriver");
    WebDriver browser = new ChromeDriver();
    try {
        browser.get("https://www.example.com");
        var pageTitle = browser.getTitle();
        System.out.println(pageTitle);
    } finally {
        browser.quit();
    }
}

You can find its source code on GitHub.

5. crawler4j

crawler4j is a library dedicated to small/medium web scraping. It supports robots.txt files.

Downsides:

  • Not-so-user-friendly visitor-based API
  • It does not handle JavaScript, limiting its effectiveness against dynamic pages
  • The project is unmaintained - the last release was in March 2018
  • Relies on a legacy com.sleepycat:je:5.0.84 artifact, which is unavailable in modern Maven repositories
  • No advanced anti-scraping measures like IP rotation, CAPTCHA solving

Here's a quick example of how to use crawler4j:

public class BasicCrawler extends WebCrawler {
    @Override
    public boolean shouldVisit(Page referringPage, WebURL url) {
        String href = url.getURL().toLowerCase();
        return href.startsWith("https://example.com/");
    }

    @Override
    public void visit(Page page) {
        String url = page.getWebURL().getURL();
        System.out.println("URL: " + url);
    }

    public static void main(String[] args) throws Exception {
        CrawlConfig config = new CrawlConfig();
        config.setCrawlStorageFolder("/data/crawl/root");
        config.setPolitenessDelay(1000);
        config.setMaxDepthOfCrawling(2);

        PageFetcher pageFetcher = new PageFetcher(config);
        RobotstxtConfig robotstxtConfig = new RobotstxtConfig();
        RobotstxtServer robotstxtServer = new RobotstxtServer(robotstxtConfig, pageFetcher);
        CrawlController controller = new CrawlController(config, pageFetcher, robotstxtServer);

        controller.addSeed("https://example.com");
        controller.start(BasicCrawler.class, 1);
    }
}

You can find its source code on GitHub.

6. Apache Nutch

Apache Nutch is an open-source web crawler and search engine software based on Apache Hadoop.

Technically, it's not a library, but a command-line tool, but it's worth mentioning due to its popularity.

Downsides:

  • It's quite complex, requires a lot of configuration, and might be an overkill for simple web scraping tasks - see our Apache Nutch tutorial
  • It does not handle JavaScript/AJAX
  • No advanced anti-scraping measures like IP rotation, CAPTCHA solving

You can find its source code on GitHub.

7. Jaunt

Jaunt is a Java library used for web scraping and data extraction from HTML and XML documents.

Jaunt is popular for its simplicity and ease of use - it was designed to hide unnecessary complexity while still providing full DOM-level control.

Here's a quick example of how to use Jaunt:

public static void main(String[] args) throws Exception {
    UserAgent userAgent = new UserAgent();
    userAgent.visit("https://example.com");
    Elements links = userAgent.doc.findEach("<a>");
    for (Element link : links) {
        System.out.println("Link: " + link.getAt("href"));
    }
}

Downsides:

  • Limited JavaScript support
  • Not available via modern Maven repositories, making it hard to integrate with modern Java projects
  • No advanced anti-scraping measures like IP rotation, CAPTCHA solving
  • The free version expires after 30 days

8. WebMagic

WebMagic is a web crawling framework that's designed to be simple and flexible. It's built on top of Apache HttpClient and HTMLUnit, providing a high-level API for web scraping.

Downsides:

  • WebMagic doesn't natively support JavaScript execution, so it might struggle with dynamic content and some anti-scraping measures.
  • No advanced anti-scraping measures like IP rotation, CAPTCHA solving

Here's a quick example of how to use WebMagic:

class WebMagicExample {

    public static void main(String[] args) {
        Spider.create(new ExampleCrawler())
          .addUrl("http://example.com")
          .thread(5)
          .run();
    }

    public static class ExampleCrawler implements PageProcessor {
        private final Site site = Site.me()
          .setRetryTimes(3)
          .setSleepTime(1000)
          .setTimeOut(10000);

        @Override
        public void process(Page page) {
            String title = page.getHtml().xpath("//title/text()").toString();
            System.out.println("Title: " + title);

            page.addTargetRequests(page.getHtml().links().regex("(https://www.example.com/\\w+)").all());
        }

        @Override
        public Site getSite() {
            return site;
        }
    }
}

You can find its source code on GitHub.

9. Gecco

Gecco is an old-school lightweight web scraping framework that's designed to be simple and flexible. It's built on top of Apache HttpClient and Jsoup, providing a high-level API for web scraping.

What sets Gecco apart is its simplicity and ease of use. It's designed to hide unnecessary complexity while still providing full DOM-level control.

It can even handle distributed crawling with the help of Redis and integrates well with Spring framework and HtmlUnit.

However, Gecco is not actively maintained, doesn't work with new Java versions, and might struggle with modern web technologies and anti-scraping measures.

Here's a quick example of how to use Gecco:

First, we need to define the model class:

@Gecco(matchUrl="https://www.example.com", pipelines="consolePipeline")
public class ExamplePage implements HtmlBean {

    private static final long serialVersionUID = 1L;

    @Text
    @HtmlField(cssPath="h1")
    private String heading;

    public String getHeading() {
        return heading;
    }

    public void setHeading(String heading) {
        this.heading = heading;
    }
}

Then, we need to define a pipeline:

package com.example.crawl;

import com.geccocrawler.gecco.pipeline.Pipeline;

public class ConsolePipeline implements Pipeline<ExamplePage> {

    @Override
    public void process(ExamplePage examplePage) {
        System.out.println("Extracted Heading:");
        System.out.println(examplePage.getHeading());
    }
}

And finally, we can run the crawler:

package com.example.crawl;

import com.geccocrawler.gecco.GeccoEngine;

public class Main {
    public static void main(String[] args) {
        GeccoEngine.create()
          .classpath("com.example.crawl") // 
          .start("https://www.example.com")
          .thread(1)
          .interval(2000)
          .run();
    }
}

Note that models and pipelines are loaded dynamically, so make sure to provide the correct package name.

You can find its source code on GitHub.

10. StormCrawler

StormCrawler is a project that provides a collection of resources for building low-latency, scalable web crawlers on Apache Storm. It's designed to be modular and scalable, making it a good choice for large-scale web crawling and search engine development.

It's worth noting that, at the time of this writing, it's still in the Apache Incubator stage, some features might not be working as expected, and the documentation is almost non-existent.

The initial configuration is quite complex, but it can be simplified by using the official Maven Archetype:

mvn archetype:generate -DarchetypeGroupId=org.apache.stormcrawler -DarchetypeArtifactId=stormcrawler-archetype -DarchetypeVersion=3.0

The part responsible for crawling is the CrawlTopology class:

public class CrawlTopology extends ConfigurableTopology {

    public static void main(String[] args) {
        ConfigurableTopology.start(new CrawlTopology(), args);
    }

    @Override
    protected int run(String[] args) {
        String[] startUrls = {"http://example.com"};
        TopologyBuilder builder = new TopologyBuilder();
        builder.setSpout("spout", new MemorySpout(startUrls));
        builder.setBolt("fetcher", new FetcherBolt()).shuffleGrouping("spout");
        builder.setBolt("parser", new JSoupParserBolt()).shuffleGrouping("fetcher");

        builder.setBolt("status", new StdOutStatusUpdater())
          .fieldsGrouping("fetcher", Constants.StatusStreamName, new Fields("url"))
          .fieldsGrouping("parser", Constants.StatusStreamName, new Fields("url"));

        getConf().put("http.agent.name", "MyCrawler");
        getConf().setDebug(true);

        try (LocalCluster cluster = new LocalCluster()) {
            cluster.submitTopology("simple-crawler", getConf(), builder.createTopology());
            Thread.sleep(30000);
        } catch (Exception e) {
            throw new RuntimeException(e);
        }

        return 0;
    }
}

You can find its source code on GitHub.

Benefits of Using Libraries

Currently, you can easily open a java.net.http.HttpClient, fetch raw HTML, and regex your way through. Even though your agent could probably do it right away, there is no good reason to build it that way. Instead of reinventing the wheel, developers can rely on tooling that makes it easy.

This is what libraries like jsoup, HtmlUnit, and Selenium do. HtmlUnit goes a step further: it's a full browser simulator, so it can execute JavaScript and follow redirects the way a real browser would. Selenium runs real browsers through WebDriver, which is the only option when a site renders its content client-side and there's no HTML to parse until the JavaScript has run.

By using Java libraries for web scraping, you'll enjoy faster development, fewer bugs, and more accurate results - all while keeping your codebase cleaner and more maintainable.

Common Challenges

Unfortunately, libraries won't solve all the challenges. Many websites employ anti-scraping measures such as CAPTCHAs, rate limiting, and IP blocking to protect their data.

The less obvious problem is JavaScript-rendered content - JavaScript-heavy websites can be particularly tricky, as much of their content is rendered dynamically and may not be accessible through simple HTTP requests.

Complex HTML structures and frequent changes to page layouts can also make extracting data frustrating.

Techniques like rotating IP addresses, user-agent spoofing, and CAPTCHA solving can make your scraper more resilient.

Conclusion

As you can see, there are plenty of Java libraries available for web scraping. The best one for you will depend on your specific use case and requirements.

  • Libraries like jsoup and HtmlUnit are excellent choices for simple tasks and static content due to their ease of use and lightweight nature.
  • Selenium is a powerful option for projects requiring full browser automation, although it can be resource-intensive and may require more setup.
  • Apache Nutch is a robust tool for large-scale web crawling and search engine development, but it may be overkill for web scraping (especially if you want to use it as a library).
  • Jaunt is actively maintained and offers a simple API for web scraping. However, it has limited JavaScript support and is not available via modern Maven repositories.
  • Gecco is a lightweight and flexible framework that's easy to use, but it's not actively maintained and may struggle with modern Java versions.
  • While Crawler4j and WebMagic offer specialized features for crawling, they lack support for modern web technologies and may require additional effort to overcome certain limitations.
  • StormCrawler is a complex but promising project, but it's still in the Apache Incubator stage.

The reason why most libraries struggle with anti-scraping measures is that it often requires a lot of resources and infrastructure to overcome them. For example, bypassing rate-limiting or CAPTCHA might require rotating IP addresses, which can be quite complex to set up. This is why platforms like ScrapingBee shine - they handle all of these challenges for you, so you can focus on scraping and not infrastructure.

Before you go, check out these related reads:

image description
Grzegorz Piwowarek

Independent consultant, blogger at 4comprehension.com, trainer, Vavr project lead - teaching distributed systems, architecture, Java, and Golang

Auto-mode picks the configuration that successfully scrapes your page

Try it now