Web Content Retrieval for Humans™

Last update: Dec 19, 2022

Overview

Lassie

https://img.shields.io/coveralls/michaelhelmick/lassie/master.svg?style=flat-square

https://img.shields.io/badge/Say%20Thanks!-:)-1EAEDB.svg?style=flat-square

Lassie is a Python library for retrieving basic content from websites.

Usage

>>> import lassie
>>> lassie.fetch('http://www.youtube.com/watch?v=dQw4w9WgXcQ')
{
    'description': u'Music video by Rick Astley performing Never Gonna Give You Up. YouTube view counts pre-VEVO: 2,573,462 (C) 1987 PWL',
    'videos': [{
        'src': u'http://www.youtube.com/v/dQw4w9WgXcQ?autohide=1&version=3',
        'height': 480,
        'type': u'application/x-shockwave-flash',
        'width': 640
    }, {
        'src': u'https://www.youtube.com/embed/dQw4w9WgXcQ',
        'height': 480,
        'width': 640
    }],
    'title': u'Rick Astley - Never Gonna Give You Up',
    'url': u'http://www.youtube.com/watch?v=dQw4w9WgXcQ',
    'keywords': [u'Rick', u'Astley', u'Sony', u'BMG', u'Music', u'UK', u'Pop'],
    'images': [{
        'src': u'http://i1.ytimg.com/vi/dQw4w9WgXcQ/hqdefault.jpg?feature=og',
        'type': u'og:image'
    }, {
        'src': u'http://i1.ytimg.com/vi/dQw4w9WgXcQ/hqdefault.jpg',
        'type': u'twitter:image'
    }, {
        'src': u'http://s.ytimg.com/yts/img/favicon-vfldLzJxy.ico',
        'type': u'favicon'
    }, {
        'src': u'http://s.ytimg.com/yts/img/favicon_32-vflWoMFGx.png',
        'type': u'favicon'
    }],
    'locale': u'en_US'
}

Install

Install Lassie via pip

$ pip install lassie

or, with easy_install

$ easy_install lassie

But, hey... that's up to you.

Documentation

Documentation can be found here: https://lassie.readthedocs.org/

Comments

Fix possible ValueError in convert_to_int caused by values like 1px

When trying to parse http://www.wired.com/wiredscience/2013/09/rim-fire-map-color-scale/ a ValueError was raised in convert_to_img, because the page has image width and height values ending in px.

I changed the function to be more liberal regarding dimension values, by extracting the digits before casting to int. I added a test for this.

Not sure though if the value should be converted to int at all or kept as a string.

opened by yaph 14

Import fails on Python3.5

It appears something is seriously broken when trying to install lassie with Python 3.5. Install goes fine but when importing I get here:

Python 3.5.0 (default, Sep 23 2015, 04:41:38)
[GCC 4.2.1 Compatible Apple LLVM 7.0.0 (clang-700.0.72)] on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>> import lassie
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/Users/ben/dev/beavy/venv/src/lassie/lassie/__init__.py", line 19, in <module>
    from .api import fetch
  File "/Users/ben/dev/beavy/venv/src/lassie/lassie/api.py", line 11, in <module>
    from .core import Lassie
  File "/Users/ben/dev/beavy/venv/src/lassie/lassie/core.py", line 13, in <module>
    from bs4 import BeautifulSoup
  File "/Users/ben/dev/beavy/venv/lib/python3.5/site-packages/bs4/__init__.py", line 30, in <module>
    from .builder import builder_registry, ParserRejectedMarkup
  File "/Users/ben/dev/beavy/venv/lib/python3.5/site-packages/bs4/builder/__init__.py", line 308, in <module>
    from . import _htmlparser
  File "/Users/ben/dev/beavy/venv/lib/python3.5/site-packages/bs4/builder/_htmlparser.py", line 7, in <module>
    from html.parser import (
ImportError: cannot import name 'HTMLParseError'

opened by gnunicorn 6

Add optional structured properties for og:image and og:video

From http://ogp.me/#structured.

The og:video tag has the identical tags as og:image.

og:image:url - Identical to og:image. og:image:secure_url - An alternate url to use if the webpage requires HTTPS. og:image:type - A MIME type for this image. og:image:width - The number of pixels wide. og:image:height - The number of pixels high.

opened by jpadilla 6
Optional support for canonical URL meta tag.

This is very roughed in, but it adds support for returning the URL as provided by the canonical link element.

There isn't anything to determine precedence with og:url.

Has passing tests, and is disabled by default.

Needed this for a project, not sure if it would be useful upstream.
enhancement

opened by jmhobbs 5
Possible relative URL in og:image

I just came accros a page with a relative path value for the og:image. Adding a call to urljoin on the src attribute in line 186 of core.py would be a possibility, but maybe it's better to check for the src prop (possibly href prop too) in _filter_meta_data and do it there. What do you think about that?

opened by yaph 5
Can't get the full article.

Hi, I want to extract the article from the source url. I got only the title of the article and small parts of it under the "description" parameter.

opened by yaseenox 4
Update requests==2.8 in setup.py, too

The changelog for the last release states, that request is now pinned at version 2.8, yet when installing the latest version of lassie, it requires (and install) version 2.6 – the setup.py hasn't been updated to reflect that change and breaks the installation. This PR corrects that.

opened by gnunicorn 4
Please allow to configure the requests session

It would be useful to be able to configure the requests session used to retrieve the requested URL.

You could perhaps initialize a default session object in the Lassie constructor, which the user could then configure, and/or add a parameter to Lassie.fetch() to override the default session.

opened by tawmas 4
Bump requests from 2.18.4 to 2.20.0
⚠️ Dependabot is rebasing this PR ⚠️

If you make any changes to it yourself then they will take precedence over the rebase.

Bumps requests from 2.18.4 to 2.20.0.

Changelog

Sourced from requests's changelog.

2.20.0 (2018-10-18)

Bugfixes

Content-Type header parsing is now case-insensitive (e.g. charset=utf8 v Charset=utf8).

Fixed exception leak where certain redirect urls would raise uncaught urllib3 exceptions.

Requests removes Authorization header from requests redirected from https to http on the same hostname. (CVE-2018-18074)

should_bypass_proxies now handles URIs without hostnames (e.g. files).

Dependencies

Requests now supports urllib3 v1.24.

Deprecations

Requests has officially stopped support for Python 2.6.

2.19.1 (2018-06-14)

Bugfixes

Fixed issue where status_codes.py's init function failed trying to append to a __doc__ value of None.

2.19.0 (2018-06-12)

Improvements

Warn user about possible slowdown when using cryptography version < 1.3.4

Check for invalid host in proxy URL, before forwarding request to adapter.

Fragments are now properly maintained across redirects. (RFC7231 7.1.2)

Removed use of cgi module to expedite library load time.

Added support for SHA-256 and SHA-512 digest auth algorithms.

Minor performance improvement to Request.content.

Migrate to using collections.abc for 3.7 compatibility.

Bugfixes

Parsing empty Link headers with parse_header_links() no longer return one bogus entry.

... (truncated)

Commits

bd84045 v2.20.0

7fd9267 remove final remnants from 2.6

6ae8a21 Add myself to AUTHORS

89ab030 Use comprehensions whenever possible

2c6a842 Merge pull request #4827 from webmaven/patch-1

30be889 CVE URLs update: www sub-subdomain no longer valid

a6cd380 Merge pull request #4765 from requests/encapsulate_urllib3_exc

bbdbcc8 wrap url parsing exceptions from urllib3's PoolManager

ff0c325 Merge pull request #4805 from jdufresne/https

b0ad249 Prefer https:// for URLs throughout project

Additional commits viewable in compare view

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.

Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

@dependabot rebase will rebase this PR

@dependabot recreate will recreate this PR, overwriting any edits that have been made to it

@dependabot merge will merge this PR after your CI passes on it

@dependabot squash and merge will squash and merge this PR after your CI passes on it

@dependabot cancel merge will cancel a previously requested merge and block automerging

@dependabot reopen will reopen this PR if it is closed

@dependabot ignore this [patch|minor|major] version will close this PR and stop Dependabot creating any more for this minor/major version (unless you reopen the PR or upgrade to it yourself)

@dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

@dependabot use these labels will set the current labels as the default for future PRs for this repo and language

@dependabot use these reviewers will set the current reviewers as the default for future PRs for this repo and language

@dependabot use these assignees will set the current assignees as the default for future PRs for this repo and language

@dependabot use this milestone will set the current milestone as the default for future PRs for this repo and language

You can disable automated security fix PRs for this repo from the Security Alerts page.

dependencies
opened by dependabot[bot] 3
Added support for open graph optional property `site_name`.

Hi, I added supported for the open graph site_name property.

This parse the following tag... <meta property="og:site_name" content="IMDb" /> into {"site_name": "IMDb"}

opened by cameronmaske 3
make image urls absolute and added mock to test_requirements

I made a change so that when lassie.fetch is called with all_images=True the images src attributes contain absolute URLs. Since lassie already comes with a function that makes relative URLs absolute, I think it's better done inside lassie than in the application which imports it.

When trying to run the tests after the changes the mock package was missing, so I added it to the test_requirements.txt file.

opened by yaph 2
docs: Fix a few typos
There are small typos in:

docs/usage/advanced_usage.rst

Fixes:

Should read attributes rather than attibutes.

Should read actual rather than acutal.

Semi-automated pull request generated by https://github.com/timgates42/meticulous/blob/master/docs/NOTE.md
opened by timgates42 0
Any reason to pindown upper version in requirements.txt
Hi,

Since lassie is a library, limiting upper versions for dependencies as in

requests>=2.18.4,<3.0.0 beautifulsoup4>=4.9.0,<4.10.0

can lead to conflicts for software using it, e.g. on pip install:

The conflict is caused by: The user requested beautifulsoup4==4.10.0 lassie 0.11.11 depends on beautifulsoup4<4.10.0 and >=4.9.0

Is there any reason for the pindown?
opened by idlesign 1
Encoding issues with german umlauts

Hi,

when getting the description from a German website the "ü" "ä" etc. end up being "Ã¤", "Ã¼" etc. Example: https://finanzguru.de/ Result:

Finanzguru - Finanzen magisch einfach Finanzen magisch einfach. Verwalte deine VertrÃ¤ge, kÃ¼ndige per Fingertipp und spare Geld mit meinen Spartipps. Alles an einem Ort und komplett kostenfrei. Einfacher war es noch nie.

I am using lassie within Django.

opened by leugh 0
Add new filters for embeddable items

The idea is to return as much data as we can in the API so users can possibly embed media. (i.e. Spotify tracks)

We'll probably add a new embed.py and return a new embed key in the lassie API response.
enhancement

opened by michaelhelmick 0

Releases(0.11.11)

0.11.11(Aug 20, 2021)
Fix PyPI release

Source code(tar.gz)
Source code(zip)
0.11.10(Aug 20, 2021)
Add html to response dict when available.

Upgrade to GitHub Actions

Source code(tar.gz)
Source code(zip)
0.11.8(Dec 16, 2020)
Update requests dependency

Source code(tar.gz)
Source code(zip)
0.11.2(Nov 1, 2017)
Add support for OEmbed providers (YouTube)

Source code(tar.gz)
Source code(zip)
0.10.0(Feb 3, 2017)
Fix issue where a website may have malformed HTML and no tag causing soup.html to be None (#60)

Updated beautifulsoup4 to 4.5.3

Update html5lib to 1.0b10

Source code(tar.gz)
Source code(zip)
0.9.0(Jan 29, 2017)
Added a default fake user agent to use instead of using python-requests/version (some websites will mark certain user agents as bot attempts)

Updated requests to 2.13.0

Source code(tar.gz)
Source code(zip)
0.8.7(Dec 21, 2016)
Fix Python 3 support

Handle empty AMP image lists

Source code(tar.gz)
Source code(zip)
0.8.6(Nov 17, 2016)
Handle AMP image list of strings vs list of objects

Source code(tar.gz)
Source code(zip)
0.8.5(Nov 3, 2016)
Handle AMP data that is contained in a list

Retrieve videos and thumbnails (as images) from AMP VideoObjects

Source code(tar.gz)
Source code(zip)
0.8.4(Nov 1, 2016)
Fix issue where AMP images could be lists inside an object

Source code(tar.gz)
Source code(zip)
0.8.3(Nov 1, 2016)
Fix issue where some keys returned (i.e. description) would not be retrieved if the key existed with an empty value already

Source code(tar.gz)
Source code(zip)
0.8.2(Nov 1, 2016)
Fix issue where AMP images could be images and not objects

Source code(tar.gz)
Source code(zip)
0.8.1(Nov 1, 2016)
Add support for AMP "description" attribute

Fix issue where an error would be thrown if width/height of an image weren't strings

Fix duplicate AMP title request, should have been url

Source code(tar.gz)
Source code(zip)
0.8.0(Nov 1, 2016)
Add support for links that use AMP

Source code(tar.gz)
Source code(zip)
0.7.1(Jul 27, 2016)
Add support for open graph site_name

Source code(tar.gz)
Source code(zip)
0.7.0(Jul 27, 2016)
Add status_code to response dictionary

Source code(tar.gz)
Source code(zip)
0.6.2(Nov 11, 2015)
Pinned requests library to version 2.8.1

Pinned beautifulsoup4 library to version 4.4.1

Add Python 3.5 to Travis CI build matrix (officially support 3.5)

Source code(tar.gz)
Source code(zip)
0.5.4(Aug 19, 2015)
Support for secure url image and videos from Open Graph

Simplified merge_settings and data updating internally

Source code(tar.gz)
Source code(zip)
0.5.3(Jul 13, 2015)

Handle when a website doesn't set a value on the "keywords" meta tag
Source code(tar.gz)
Source code(zip)
0.3.0(Aug 17, 2013)
Added support for locale to be returned. If lang is specified in the html tag and it normalizes to an actual locale, it will be added to the returned data.

Fixed bug where height was not being returned for body images

Added test coverage, we're 100% covered! :D

Source code(tar.gz)
Source code(zip)
0.2.0(Aug 6, 2013)

Fix package error when importing
Source code(tar.gz)
Source code(zip)
0.1.0(Aug 6, 2013)

Initial Release
Source code(tar.gz)
Source code(zip)

Owner

Mike Helmick

GitHub Repository https://lassie.readthedocs.org

A database scraper created with mechanical soup and sqlite

WebscrapingDatabases a database scraper created with mechanical soup and sqlite author: Mariya Sha Watch on YouTube: This repository was created to su

30 Aug 08, 2022

PaperRobot: a paper crawler that can quickly download numerous papers, facilitating paper studying and management

PaperRobot PaperRobot 是一个论文抓取工具，可以快速批量下载大量论文，方便后期进行持续的论文管理与学习。 PaperRobot通过多个接口抓取论文，目前抓取成功率维持在90%以上。通过配置Config文件，可以抓取任意计算机领域相关会议的论文。 Installation Down

47 Nov 23, 2022

Web Crawlers for Data Labelling of Malicious Domain Detection & IP Reputation Evaluation

Web Crawlers for Data Labelling of Malicious Domain Detection & IP Reputation Evaluation This repository provides two web crawlers to label domain nam

1 Nov 05, 2021

Works very well and you can ask for the type of image you want the scrapper to collect.

Works very well and you can ask for the type of image you want the scrapper to collect. Also follows a specific urls path depending on keyword selection.

1 Feb 17, 2022

download NCERT books using scrapy

download_ncert_books download NCERT books using scrapy Downloading Books: You can either use the spider by cloning this repo and following the instruc

1 Dec 02, 2022

Amazon web scraping using Scrapy Framework

Amazon-web-scraping-using-Scrapy-Framework Scrapy Scrapy is an application framework for crawling web sites and extracting structured data which can b

1 Jan 25, 2022

Haphazard scripts for scraping bitcoin/bitcoin data from GitHub

This is a quick-and-dirty tool used to scrape bitcoin/bitcoin pull request and commentary data. Each output/pr number folder contains comments.json:

8 Oct 12, 2022

script to scrape direct download links (ddls) from google drive index.

bhadoo Google Personal/Shared Drive Index scraper. A small script to scrape direct download links (ddls) of downloadable files from bhadoo google driv

53 Dec 16, 2022

12306抢票脚本

457 Jan 05, 2023

A module for CME that spiders hashes across the domain with a given hash.

hash_spider A module for CME that spiders hashes across the domain with a given hash. Installation Simply copy hash_spider.py to your CME module folde

37 Sep 08, 2022

Examine.com supplement research scraper!

ExamineScraper Examine.com supplement research scraper! Why I want to be able to search pages for a specific term. For example, I want to be able to s

15 Dec 06, 2022

Iptvcrawl - A scrapy project for crawl IPTV playlist

iptvcrawl a scrapy project for crawl IPTV playlist. Dependency Python3 pip insta

18 May 05, 2022

A dead simple crawler to get books information from Douban.

Introduction A dead simple crawler to get books information from Douban. Pre-requesites Python 3 Install dependencies from requirements.txt (Optional)

1 Jan 10, 2022

Python script for crawling ResearchGate.net papers✨⭐️📎

ResearchGate Crawler Python script for crawling ResearchGate.net papers About the script This code start crawling process by urls in start.txt and giv

4 Aug 30, 2022

API which uses discord to scrape NameMC searches/droptime/dropping status of minecraft names

NameMC Scrape API This is an api to scrape NameMC using message previews generated by discord. NameMC makes it a pain to scrape their website, but som

2 Dec 22, 2021

NASA APOD Discord Bot - Fetches information from NASA APOD site.

4 Apr 23, 2022

一个m3u8视频流下载脚本

一个Python的m3u8流视频下载脚本介绍 m3u8流视频日益常见，目前好用的下载器也有很多，我把之前自己写的一个小脚本分享出来，供广大网友使用。写此程序的目的在于给视频下载爱好者提供一个下载样例，可直接调用，勿再重复造轮子。使用方法在python中直接运行程序或进行外部调用 import

0 Oct 10, 2021

A leetcode scraper to compile all questions in leetcode free tier to text file. pdf also available.

A leetcode scraper to compile all questions in leetcode free tier to text file, pdf also available. if new questions get added, run again to get new questions.

3 Dec 07, 2021

Comment Webpage Screenshot is a GitHub Action that captures screenshots of web pages and HTML files located in the repository

Comment Webpage Screenshot is a GitHub Action that helps maintainers visually review HTML file changes introduced on a Pull Request by adding comments with the screenshots of the latest HTML file cha

21 Sep 29, 2022

A scrapy pipeline that provides an easy way to store files and images using various folder structures.

scrapy-folder-tree This is a scrapy pipeline that provides an easy way to store files and images using various folder structures. Supported folder str

7 Oct 23, 2022

Web Content Retrieval for Humans™

Related tags

Overview

Lassie

Usage

Install

Documentation

Comments

2.20.0 (2018-10-18)

2.19.1 (2018-06-14)

2.19.0 (2018-06-12)

Releases(0.11.11)

0.11.11(Aug 20, 2021)

0.11.10(Aug 20, 2021)

0.11.8(Dec 16, 2020)

0.11.2(Nov 1, 2017)

0.10.0(Feb 3, 2017)

0.9.0(Jan 29, 2017)

0.8.7(Dec 21, 2016)

0.8.6(Nov 17, 2016)

0.8.5(Nov 3, 2016)

0.8.4(Nov 1, 2016)

0.8.3(Nov 1, 2016)

0.8.2(Nov 1, 2016)

0.8.1(Nov 1, 2016)

0.8.0(Nov 1, 2016)

0.7.1(Jul 27, 2016)

0.7.0(Jul 27, 2016)

0.6.2(Nov 11, 2015)

0.5.4(Aug 19, 2015)

0.5.3(Jul 13, 2015)

0.3.0(Aug 17, 2013)

0.2.0(Aug 6, 2013)

0.1.0(Aug 6, 2013)