How can I prevent the new indexing bug from creating useless pages on Google?

@gregbernhardt what are we supposed to do? Wait for you to fix this at Shopify or am I best employing an SEO expert to fix it on our end?

I don’t want to have issues if the SEO expert fixes it and then you implement a fix which causes more problems with clashing fixes.

Can you give me a straight answer, do I:

  1. Leave it and Shopify will fix this for all merchants across the board

OR

  1. Employ someone to fix this issue on my website/theme alone

We can’t leave it like this.

I was the 2nd person to report this issue in this post well back in 2022 , I can’t believe that it’s taking so long to not even get a fix but just get answers that make some kind of sense or some legitimate assurances other than a document will follow soon with no specific timeline mentioned.

My impression is going bottom and I still don’t have any idea to fix it. Since middle of march, my organic traffic has significantly dropped. Please advise me if u guys have any idea to solve this.

No Index Tag: 700> 11,000 since Mid Mar

Canonical tag: 5000> 13000 since Mid Mar

@George_Greenhil you don’t need to hire anyone. Please wait for our official document to be released very soon.

Hello.

Thank you to everyone who has shared their feedback. We’re currently investigating this issue and are working on a FAQ that will help address the most common concerns. We will update you here once we have this FAQ ready.

I will be marking my reply as the solution as this is the easiest way to surface the most up to date information on the current status of this thread. I will remove the solution once the FAQ is ready.

Thank you.

Hi Trevor, looking forward to the update. Can we expect to see a solution, or should we expect to only see an explanation of the events that transpired? The words “FAQ” and “help address” aren’t inspiring confidence that this is going to be an issue that Shopify fixes on Shopify’s side using Shopify’s time and resources.

So far I’ve seen a lot of blame on Google which is weird because our traffic dropped before the core update. Our first major traffic drop was the 13th into the 14th. The google core update rollout was officially started between the 15th and the 22nd, finishing on the 28th.

I know you guys are doing what you can, but the sooner we get guidance the better. I have a client who has pretty much demanded I submit almost 50,000 301 redirects for them today, one for every page and every single version of WPM that was released, they’re desperate and getting ready to furlough people due to the reduced traffic. I’m having to play defense for Shopify here and so far it has not been easy.

I’ve watched this client make a ton of sales on Shopify. They’re one of my top clients as well. Lets get this fixed so we both don’t lose an excellent customer.

Hi Kbarcant, can we connect? We both kind of having a similar situation and our google search console graph basically show the same pattern. I wish we can both sort this out. I will private message you in a short while. :slightly_smiling_face:

We have same situation with drop but without “not indexed” increase(enormous “not indexed” was bit earlier) My guess because google was ignoring all params url in general. Right now they change logic and somehow make intact in SEO.

If everyone can please DM me a screenshot of your GSC noindex report along with the impression line as @DariusWS and @kbarcant did above that would be a great help, thanks!

Wow, things have really changed since I last looked at this! On 4-18-23 I was a t 2.37M not indexed 21.6K indexed. Today, 4-26-23, I’m at 961K indexed and 6.76 indexed. I haven’t seen these low numbers since before the exploits began for me in late November around Thanksgiving.

I do still have the following code in robots.txt but no removals in GSC:

{%- if request.page_type == ‘search’ and search.performed and search.results_count == 0 -%} {%- endif -%}

Hi,
I don’t know how to DM you, but here’s an overview of one of my sites.

Besides the enormous increase of the “web-pixels” and “wpm” URLs found in GSC, my index gets bloated by a large number of dynamic URLs coming out of “Product recommendation” switched on. Google sorted these out in the past, now these days Google started to index them and created many indexed content duplicates. I have to switch this functionality off as it does more harm than good.

We had the same problem longer ago. The URL strcuture was:

/recommendations/products?product_id=6963005XXXXXX&limit=4&section_id=template–16191536XXXXXX__product-recommendations

As far as i remember, those URLs were not indexed, they were only crawled. Anyway, it is better to exclude these from crawling and to optimize crawling budget. To solve this, you need to add a robots.txt statement.

In our case it was neccesary to add the following robots.txt statement:

Disallow: /*section_id=template--*

This was part of the URL with the product-recommendations. There are also different other robots.txt statements you could exclude those links with, e.g.:

Disallow: /recommendations/*product-recommendations*

To test, if the statement is working, you can use the robots.txt Testing Tool: https://support.google.com/webmasters/answer/6062598?hl=en

2 Things to concider:

  • Make sure, that there are no important URLs you exclude from crawling with this statement (e. g. by using Google Analytics, looking for URLs with the same structure, which are important to you)
  • Make sure that none of these URLs are indexed as robots.txt won’t kick those URLs out of the index. In this case, you’d first need to put a “noindex”-Tag to those pages and wait for Google to kick all of them out of the index. After that you can add the statement to the robots.txt. As far as I remember, those pages already have the noindex-Tag, so that none of these URLs shall be indexed.

We are continuing to see a rise in “not indexed” pages with every crawl.

404s look to have peaked in early April however, we keep seeing an increase in “Excluded by ‘noindex’ tag” and “Crawled - currently not indexed” in GSC. Of note, the last crawl saw a big spike in “Crawled - currently not indexed” for the [email removed] pages.

We can also confirm the same pattern many store owners are reporting of a sharp decline in search impressions reported by GSC starting in mid-March.

  • Is there anything for store owners to do at this juncture?
  • Is this happening to EVERY Shopify store?
  • Can Shopify interface with Google and let them know to ignore the @WPM pages while a fix is being implemented?

Wait for the document. Do not use the temporary URL removal tool in the mean time, it does not have an undo function.

Hi @romko18

Can you give us more information on the large number of dynamic urls coming out of “product recommedation” . I havent noticed anything like that and if that is effecting my GSC indexing i would have to turn it off too!

IMO Shopify nor Google will solve this… no worries, their prerogative, but we can’t wait any longer.

Yes, we can’t wait longer, my store has a huge impact on this. I simply just did nothing and it just suddenly hits me.

Hi @shadi1 ,

the dynamic product recommendation URLs contain the same pattern: “prod_strat”.
In the past, they were “Discovered, but not indexed”. Some weeks, or months ago I’ve seen that Google started to consider them more, I found some in 404, and many were in the “Discovered” section. This was telling me I have to exclude crawling them in Robots.txt.

But this didn’t obviously help to prevent indexing them. To set “noindex” in Shopify is not so easy and in my case, it didn’t help:

Thanks @romko18 Another issue to look out for. I have been noticing false reporting or phantom sessions on shopify live traffic reporting. Its reporting alot more traffic then what google analytic live view is. Not sure whats going on but maybe its related to what shopify team is doing to fix this issue?

Any one experiencing this on there shopify live traffic reporting?

@Mont everyone is at the end of their rope. I’ve pretty much held off my clients for the weekend but I’m seeing big changes coming in my life if Shopify doesn’t fix this or get Google to intervene somehow. Unfortunately some of the smaller stores I talk to will likely be better off with new domains, the small traffic they were getting is now near zero so why not? Why build on a foundation of trash? This whole situation makes me sick. It’s a shame Gilbert Gottfried is gone, I would’ve paid to have him read some of the customer service we’ve gotten, maybe make this situation a little more light-hearted.

Google is still discovering new WPM pages on their site too, shopify isn’t even doing a good job hiding them. On 4/23 google crawled 17 new [email removed] pages for one of my clients. Google shouldn’t be wasting our crawl budget on these. The crawler shouldn’t be able to see them at all.

About the Google Temporary Removal Tool

Unfortunately I’ve gotten some messages from people about to use this tool not understanding what it does. Do not use this tool unless you are 100% sure you understand how it works, what it does, and what it can do to your website. In my opinion, there’s a non-zero chance of apocalyptic results. This situation also sounds like it firmly belongs under the “When Not To Use This Tool” category on their documentation Source.

Why The Temporary Removal Tool Might Be An Awful Idea

  • This tool removes a URL from the search results. Lets say we want to hide a specific product, we’d submit it as /products/productname. The thing is, this product is also viewable on another page, a “non-canonical” page, such as /collections/collectionname/products/productname. That product has a tag in it that tells Google it’s the non-canonical version. All URLs associated through canonicalization are removed Source

  • Canonicalization was likely declared in many older versions of these blank pages as it’s rendered in theme.liquid.

  • What link did it declare? Did {{canonical_url}} return a liquid error? Did it return blank? Does that mean it canonicalized them as our home page? Or worse, did it submit the regular URL for that page? Unfortunately, we have no way to check. We have to assume these pages reported valid URLs to google as the canonicalized version of the page.

  • The file in the google index might be the older version with canonical URLs still in the google indexed version. With these thousands, or even millions of bad pages, how are we to verify google doesn’t have a version with a poison pill in hand?

  • When we submit the removal, Google likely won’t re-crawl for the latest version. I don’t see any mention of that in the documentation. This means they probably use their last indexed version - if they didn’t have an indexed version, you wouldn’t be asking for removal.

  • I see nothing about ignoring canonicalization removals for wildcard removed pages. What I do see is advice saying not to use this to remove one version of a page that’s canonicalized as it will remove all versions.

  • Lets pretend they fetch the lastest page

    • A fetch of the latest version has two options. One option is a 404, and google gets no information on canonicalization. Most are going to result in a 404. Do they just write-off the canonicalization link between the two pages now?
    • The other option is they see the actual, real life WPM pages that are still existing on our websites (Shopify, please delete these, your web pixels manager isn’t worth it, we don’t want it, we were better off before it). These pages do not have any canonicalization declared, so does google just forget the old link between that page and the other?

In my opinion this tool should not be utilized as a blanket solution.

@shadi1 I had 88 people from Singapore yesterday before lunch, for a second I thought my traffic was picking back up.