# Get HTML of remote page with JS

**URL:** <https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709>\
**Category:** Development\
**Created:** [April 15, 2018, 2:03pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709 "2018-04-15T14:03:37Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 15, 2018, 2:03pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/1 "2018-04-15T14:03:37Z")

</div>

How can I get HTML content of remote page by given url? In particular I need it’s title and whole body.

---

<div class="post-metadata">

**Author:** ![NilkasG](https://avatars.discourse-cdn.com/v4/letter/n/0ea827/32.png) [@NilkasG](https://discourse.mozilla.org/u/NilkasG)\
**Post date:** [April 15, 2018, 8:15pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/2 "2018-04-15T20:15:33Z")

</div>

```auto
const html = (await (await fetch(url)).text()); // html as text
const doc = new DOMParser().parseFromString(html, 'text/html');
doc.title; doc.body;

```

---

<div class="post-metadata">

**Author:** ![freaktechnik](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/freaktechnik/32/37766_2.png) [@freaktechnik](https://discourse.mozilla.org/u/freaktechnik)\
**Post date:** [April 15, 2018, 8:26pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/3 "2018-04-15T20:26:20Z")

</div>

You probably meant `.text()`

Also `view` is the global scope, I assume (i.e. `window`).

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 15, 2018, 9:09pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/4 "2018-04-15T21:09:54Z")

</div>

I tried this and browser said request has been blocked because CORS herader is missing. And then

> TypeError: NetworkError when attempting to fetch resource.

(Because fetch was failed I guess?)

---

<div class="post-metadata">

**Author:** ![freaktechnik](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/freaktechnik/32/37766_2.png) [@freaktechnik](https://discourse.mozilla.org/u/freaktechnik)\
**Post date:** [April 15, 2018, 10:04pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/5 "2018-04-15T22:04:38Z")

</div>

You do need a host permission for the page in most cases to circumvent CORS restrictions, yes.

---

<div class="post-metadata">

**Author:** ![NilkasG](https://avatars.discourse-cdn.com/v4/letter/n/0ea827/32.png) [@NilkasG](https://discourse.mozilla.org/u/NilkasG)\
**Post date:** [April 15, 2018, 11:18pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/6 "2018-04-15T23:18:17Z")

</div>

Yes and yes. I typed the first part without my glasses and copied the second one. For reference I corrected it above.

Regarding host permission:  
Yes, one does indeed need it (`<all_urls>` is the easiest way to test it). Here is why:

- [https://developers.google.com/web/ilt/pwa/working-with-the-fetch-api#cross-origin\_requests](https://developers.google.com/web/ilt/pwa/working-with-the-fetch-api#cross-origin_requests)
- having a host permission for the target drops that cross-origin restriction (and CORS would actually be another way around it, if it is supported by the server)

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 16, 2018, 6:25am UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/7 "2018-04-16T06:25:40Z")

</div>

Is there answer by your link or I read bad? So what should I do? I saw extensions somehow did get remote HTML.

---

<div class="post-metadata">

**Author:** ![NilkasG](https://avatars.discourse-cdn.com/v4/letter/n/0ea827/32.png) [@NilkasG](https://discourse.mozilla.org/u/NilkasG)\
**Post date:** [April 16, 2018, 6:35am UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/8 "2018-04-16T06:35:14Z")

</div>

The two bullet points after “Here is why” just explain _why_ you need a host permission.

I’m not sure what you mean wit the rest of your comment:

> I saw extensions somehow did get remote HTML.

So it works? Great!

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 16, 2018, 6:40am UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/9 "2018-04-16T06:40:54Z")

</div>

> [@NilkasG](#):
>
> I’m not sure what you mean wit the rest of your comment

I meant somewhose other’s extension did it: [Group Speed Dial](https://addons.mozilla.org/en-US/firefox/addon/groupspeeddial/) can get page title and take a screenshot of it by user’s url. So it can be done, but question is how.

---

<div class="post-metadata">

**Author:** ![NilkasG](https://avatars.discourse-cdn.com/v4/letter/n/0ea827/32.png) [@NilkasG](https://discourse.mozilla.org/u/NilkasG)\
**Post date:** [April 16, 2018, 3:14pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/10 "2018-04-16T15:14:35Z")

</div>

The code to load the **source** of a page and get its title and body (as DOM element) from that is in my first comment.

So far, you didn’t say that you want a _screenshot_ of the rendered page. That is only possible to get from open/loaded pages. I am quite sure that the linked extension just waits for the page to be loaded by the user and grabs the screenshot then.

An alternative would be to render the pages on a server (this can work with `puppeteer`).

With tab hiding (experimental in Firefox) you may also be able to just load the desired url in a hidden tab and take a screenshot there. Loading hidden tabs may have unforeseen consequences, though.

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 16, 2018, 3:32pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/12 "2018-04-16T15:32:24Z")

</div>

I found lib that generates screenshot by DOM element so discussion is relevant.  
I didn’t hear about hidden tabs, it may be interesting. As for the linked extension, there are two options to take a screenshot: just take and take via visiting.

I will try a trick with hidden tab a bit later. My first idea is to read all needed info with content script and send it to main script (wondering how to). Or it can be done easier?

---

<div class="post-metadata">

**Author:** ![NilkasG](https://avatars.discourse-cdn.com/v4/letter/n/0ea827/32.png) [@NilkasG](https://discourse.mozilla.org/u/NilkasG)\
**Post date:** [April 16, 2018, 5:14pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/13 "2018-04-16T17:14:40Z")

</div>

It really depends on what exactly you intend to achieve. If the page is already loaded, [tabs.captureTab() - Mozilla | MDN](https://developer.mozilla.org/en-US/Add-ons/WebExtensions/API/tabs/captureTab) seems the most straight-forward solution.

> My first idea is to read all needed info with content script and send it to main script (wondering how to).

I don’t know what “needed info” you refer to and how any information (except for the entire DOM serialized with evaluated inline styles) could ever let you render an external page in a background script.

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 16, 2018, 6:36pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/14 "2018-04-16T18:36:46Z")

</div>

> [@NilkasG](#):
>
> I don’t know what “needed info” you refer to

Document title and body as I said

---

<div class="post-metadata">

**Author:** ![NilkasG](https://avatars.discourse-cdn.com/v4/letter/n/0ea827/32.png) [@NilkasG](https://discourse.mozilla.org/u/NilkasG)\
**Post date:** [April 16, 2018, 8:56pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/15 "2018-04-16T20:56:47Z")

</div>

Document title is pretty clear, it’s a string, but “document body” **_in what form_**?

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 16, 2018, 9:27pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/16 "2018-04-16T21:27:58Z")

</div>

In form that can be rendered to image. DOM element is suitable at the moment unless my idea about getting it from website and sending from content script to script inserted into my own HTML page is too hard to implement.

---

<div class="post-metadata">

**Author:** ![NilkasG](https://avatars.discourse-cdn.com/v4/letter/n/0ea827/32.png) [@NilkasG](https://discourse.mozilla.org/u/NilkasG)\
**Post date:** [April 16, 2018, 9:37pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/17 "2018-04-16T21:37:13Z")

</div>

Ok. If your goal is to render arbitrary page bodies in the background the same way they are / would be rendered on the webpage, then that can’t be done (or is very difficult). The reason is mostly that modern web pages do not only consist of HTML. When you fetch or serialize the body element as a string, you are missing information. And reconstructing that in general, without actually executing and rendering the entire page, it very far from trivial.

So (as I said):  
You need to take (maybe partial) screenshots of the actual running page. As I said, that can’t be done in the background. You can either do it in a browser tab (maybe already open, maybe hidden) or on a server.

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 16, 2018, 9:43pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/18 "2018-04-16T21:43:42Z")

</div>

Sad to know. Your method is promising though. I’ll try it one of these days and inform about successes.

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 20, 2018, 1:29pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/19 "2018-04-20T13:29:28Z")

</div>

> [@NilkasG](#):
>
> It really depends on what exactly you intend to achieve. If the page is already loaded, [tabs.captureTab() - Mozilla | MDN](https://developer.mozilla.org/en-US/Add-ons/WebExtensions/API/tabs/captureTab) seems the most straight-forward solution.

Thanks, it works and it is much easier than I imagined:

```
browser.tabs.create({url: "https://developer.mozilla.org/", active: false}).then(function (tab) {
    console.log("Tab:", tab);
    setTimeout(function () {
        browser.tabs.captureTab(tab.id).then(function(base64img) {
            console.log("Title:", tab.title); 
            console.log("Favicon:", tab.favIconUrl);
            console.log("Base64img", base64img);
            browser.tabs.remove(tab.id);
        });
    }, 3000); //give page some time to load itself
});

```

But I wonder why `tab.title` is just url and `tab.favIconUrl` is undefined. Extension has `tabs` and `<all_urls>` permissions.

---

<div class="post-metadata">

**Author:** ![freaktechnik](https://sea1.discourse-cdn.com/flex001/user_avatar/discourse.mozilla.org/freaktechnik/32/37766_2.png) [@freaktechnik](https://discourse.mozilla.org/u/freaktechnik)\
**Post date:** [April 20, 2018, 1:31pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/20 "2018-04-20T13:31:54Z")

</div>

Probably because you’re not actually waiting for the tab to load and instead just wait an arbitrary number of seconds.

---

<div class="post-metadata">

**Author:** ![MeowHellYeah](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@MeowHellYeah](https://discourse.mozilla.org/u/MeowHellYeah)\
**Post date:** [April 20, 2018, 1:38pm UTC](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709/21 "2018-04-20T13:38:00Z")

</div>

True, I am, but it’s enough to load page completely and take a good screenshot.  
Btw I don’t see anything like `tabs.onLoaded` event in `tabs` API. What is the good way?

[Next page](https://discourse.mozilla.org/t/get-html-of-remote-page-with-js/27709.md?page=2)
