XPath Scraping with FreshRSS

I’ve been spending a while running on reduced brain capacity lately so, to ease myself back into thinking like a programmer, I upgraded my preferred feed reader FreshRSS to version 1.20.0 – which was released a couple of weeks ago – and tried out what I believe is its killer new feature: HTML + XPath scraping.

Screenshot showing Beverley Newing's weblog; two articles are visible - Paperback copy of 'Disability Visibility', edited by Alice Wong, next to a cup of tea Setting up an Accessibility Book Club, published on 1 March 2022, and Reflecting on 2021, published on 1 January 2022. — I like to keep up-to-date with my friend Bev’s blog, but they don’t have an RSS feed.

I’ve been using RSS¹ for about 20 years and I love it. It feels great to be able to curate my updates based on “what I care about”, and not on “what some social network thinks I should care about”, to keep things to read later, to prioritise effectively based on my own categorisation, to consume content offline and have my to-read list synchronise later, etc.

RSS never went away, of course (what do you think a podcast is?), but it got steamrollered out of the public eye by big companies who make their money out of keeping your eyes on their platforms and off the open Web. But it feels like it’s slowly coming back: even Substack – whose entire thing is that an email client is more-convenient than a feed reader for most people – launched an RSS reader this week!

I love RSS so much that I routinely retrofit other people’s websites with feeds just so I can subscribe to them: I even published the tool I use to do so! Whether filtering sports headlines out of BBC News, turning retro webcomics into “reading lists” so I can track my progress, or just working around sites that really should have feeds but refuse to, I just love sidestepping these “missing feeds”. My friend Beverley has a blog without any kind of feed, so I added one so I could subscribe to it. Magic.

But with FreshRSS 1.20.0, I no longer have to maintain my own tool to get this brilliant functionality, and I’m overjoyed. Let’s look at how it works by re-subscribing to Beverley’s blog but without a middleware tool.

In the latest version of FreshRSS, when you add a new feed to your reader, a new section “Type of feed source” is available. Unfold it, and you can change from the default (“RSS / Atom”) to the new option “HTML + XPath (Web scraping)”. Put a human-readable page address rather than a feed address into the “Feed URL” field and fill these fields to tell FreshRSS how to parse the page to get the content you want. Note that it doesn’t matter if the web page isn’t valid XML (e.g. missing closing tags) because it’s going to get run through PHP’s DOMDocument anyway which will “correct” for some really sloppy code if needed.

You’ll need to use XPath to express how to find a “feed item” on the page. Here’s the rules I used for https://webdevbev.co.uk/blog.html (many of these fields were optional – I didn’t have to do this much work):

Feed title: //h1
I override this anyway in FreshRSS, so I could just have used the a string, but I wanted the XPath practice. There’s only one <h1> on the page, and it can be considered the “title” of the feed.
Finding items: //li[@class="blog__post-preview"]
Each “post” on the page is an <li class="blog__post-preview">.
Item titles: descendant::h2
Each post has a <h2> which is the post title. The descendant:: selector scopes the search to each post as found above.
Item content: descendant::p[3]
Beverley’s static site generator template puts the post summary in the third paragraph of the <li>, which we can select like this.
Item link: descendant::h2/a/@href
This expects a URL, so we need the /@href to make sure we get the value of the <h2><a href="...">, rather than its contents.
Item thumbnail: descendant::img[@class="blog__image--preview"]/@src
Again, this expects a URL, which we get from the <img src="...">.
Item author: "Beverley Newing"
Beverley’s blog doesn’t host any guest posts, so I just use a string literal here.
Item date: substring-after(descendant::p[@class="blog__date-posted"], "Date posted: ")
This is the only complicated one: the published dates on Beverley’s blog aren’t explicitly marked-up, but part of a string that begins with the words “Date posted: “, so I use XPath’s substring-after function to strtip this. The result gets passed to PHP’s strtotime(), which is pretty tolerant of different date formats (although not of the words “Date posted:” it turns out!).

Screenshot: Adding a "HTML + XPath (Web scraping)" feed via FreshRSS. — I’d love one day for FreshRSS to provide some kind of “preview” feature here so you can see what you’ll expect to get back, as you work. That, and support for different input types (JSON, perhaps?), perhaps other selectors (I find CSS-style selectors much simpler than XPath), and maybe even an option to execute Javascript on the page before scraping (I use this in my own toolchain, but that’s just because I want to have my cake and eat it too). But this is still all pretty awesome.

I hope that this is just the beginning for this new killer feature in FreshRSS: there’s so much more it can be and do. But for now, I’m still mighty impressed that I can begin to phase-out my use of my relatively resource-intensive feed-building middleware and use my feed reader to do more and more of the heavy lifting for which I love it so much.

I also love that this functionally adds h-feed support in by the back door. I’d still prefer there to be a “h-feed” option in the “Type of feed source” drop-down, but at least I can add such support manually, now!

Beverley's blog post "Setting up an Accessibility Book Club" in FreshRSS. — The finished result: Bev’s blog posts appear directly in my feed reader, even though they don’t have a feed, and now *without* going through the middleware I’d set up for that purpose.

Footnotes

¹ When I say RSS, I mean feed. Most of the feeds I subscribe to are RSS feeds, but some are Atom feeds, h-feed, etc. But I can’t get over the old-fashioned name, and I don’t care to try.

12 comments

FreshRSS says:

Really great article, thanks! 😍

Read more →

27 September, 2022, 18:01
Alkarex says:

Thanks for the great article 👍🏻
I would love to get your feedback on our pull requests as well as FreshRSS release candidates, so do not hesitate to reach out! https://github.com/FreshRSS/FreshRSS/
(Also if you write other FreshRSS articles – some could even be linked from our documentation – PRs welcome)
More precise ideas regarding h-card and JSON are also welcome (I have been thinking about options for JSON) already, in particular regarding how often those use-cases could be used on Web sites not also providing RSS/ATOM feeds.

1 October, 2022, 16:58
David says:

Thank you. Can you suggest a reader? Im hooked on Feedly because I’ve used it so long and gotten used to swiping left and right, and tagging things. Other readers feel weird. I will use freshrss too self hosted, but what reader apps do you think are the best ones? Im on Mac and iOS. Cheers!

26 December, 2023, 20:43
Dan Q says:

I use FreshRSS’s own Web-based reader on desktop. It’s fast, responsive, always up-to-date, has sensible keyboard shortcuts, and benefits from some FreshRSS-specific functionality like custom JS (with a standard plugin): I use this to eg swap out the low-res images in one particular feed with the high-res variants they hide in a custom property, and use one on xkcd to show the title text below the image.

On my phone, I use FeedMe (for Android, available on f-Droid, free/donation-supported), which gives me a solid offline sync so I can read a load of news and blogs while on aeroplanes and have then marked read on FreshRSS when I touch down. Otherwise I don’t use an app at all these days! FreshRSS’s Web interface is perfectly good for my needs; I keep it in a Firefox “Pinned Tab” so it’s always handy.

26 December, 2023, 21:51
Alkarex says:

JSON support is about to land https://github.com/FreshRSS/FreshRSS/pull/5662
Feedback welcome!

3 January, 2024, 14:50
Sascha says:

Hi Dan,
I keep using the outdated rawdog, see https://offog.org/code/rawdog/ running at a virtual machine that gets no updates to keep it running with Python 2.7
Nice is, it is configure file controlled and creates one html page out of a template which I have enriched. I like that linux cli style :-)

Can your FreshRSS do the same? Create a html file filled with fetched RSSs created from a user configurable template?

11 January, 2024, 13:15
1. Dan Q says:
  
  Not as far as I know. FreshRSS is a feed reader (which happens to be able to do XPath scraping) rather than a general-purpose format converter.
  
  11 January, 2024, 16:42
setop says:

You could build a scraping repository/referential where these xpath entered by a user for a website are proposed to other users wanting to scrap the same website.

29 January, 2024, 22:31
1. Dan Q says:
  
  I love that idea!
  
  29 January, 2024, 22:41
David Whelan says:

This was really helpful. I hadn’t realized FreshRSS had this feature, even though I’ve been using it awhile. But I was noticing RSS feeds go missing and wanted an alternative to the scraper apps (only because of the cost). These instructions were really clear and I’ve been able to add a couple of feeds that were otherwise RSS-free. Thanks for sharing.

5 October, 2024, 19:43
Nick says:

Thank you for the good article.
Is there a way to debug the scraping results? I always get an error
“Blast! This feed has encountered a problem. Please verify that it is always reachable then update it.”
Possibly the scraping is forbidden, but I could also made an other mistake.

5 November, 2024, 23:17
1. Dan Q says:
  
  I’m not sure what’s behind that message, but I’ve blogged a walkthrough of my process of setting up a feed for another site. Maybe that’ll help?
  
  It’s possible that they’re attempting to block your scraper by its user agent, yes, but that’s not common. You might be able to work around it by hacking about in your FreshRSS source to fake a different user agent I suppose, if so…
  
  8 November, 2024, 19:04