Introducing the Adweek Podcast Network. Access infinite inspiration in your pocket on everything from career advice and creativity to metaverse marketing and more. Browse all podcasts.
TikTok is the latest tech platform facing regulatory heat around allegations of its data sharing. While those are sensitive claims for the U.S. administration—and user privacy has become a growing concern—history has shown it will have a limited negative impact on advertisers’ spending on the platform.
“At this point for most advertisers, the audience on TikTok is just too good to resist,” Insider analyst Jasmine Enberg told Adweek.
While other platforms are seeing declines in time spent, 2024 projections show that U.S. adults will spend nearly 20% of their social media time on TikTok. The platform, owned by Chinese tech company ByteDance, is becoming an increasingly vital tool in the performance marketer’s arsenal. Meanwhile, TikTok’s U.S. ad revenue is expected to grow to $6.83 billion this year, from $5.03 billion in 2022, per Insider Intelligence. To that, its users are projected to grow to 102.4 million in 2023. Last year, TikTok saw a total of 95.8 million users.
Still, U.S. lawmakers want to hear from the company. In March, CEO Shou Zi Chew will go before the House Energy and Commerce Committee to testify on the video-sharing app’s privacy and security practices and the concern the Chinese government could access the data of millions of users.
“TikTok has knowingly allowed the ability for the Chinese Communist Party to access American user data,” Rep. Cathy McMorris Rodgers, chair of the House Energy and Commerce Committee, said in a written statement.
Regulators are getting more aggressive with their scrutiny of Big Tech this year. TikTok has been in yearslong negotiations with the Committee on Foreign Investment (CIFUS) over its security measures. The app recently attempted to increase its transparency protocols to placate lawmakers while drawing the attention of some marketers. Although a growing number of states have banned the download of TikTok on government devices, the chair of the Senate Intelligence Committee, Sen. Mark Warner, is considering a bill that goes beyond those measures.
Not immune to the weak economic climate
Concerns of an economic downturn loom large. Market analysts anticipate struggles with the ad economy ahead of this week’s Big Tech earning calls.
This month, a survey of 50 ad buyers, who collectively spend about $23 billion annually by investment firm Cowen,showed that companies expect a slight increase in ad budgets of just 3% year over year, down from 7.5% in 2022.
Insider Intelligence lowered its TikTok ad spend forecast for 2023 by $1.92 billion. In March, 2022, the market research firm predicted TikTok’s U.S ad revenue to be $8.75 billion. That number is now at $6.83 billion.
“There are a few ad dollars to go around, and TikTok isn’t immune to that challenge,” said Enberg.
While it’s unclear what will take place within Chew’s congressional testimony, it will be another consideration for marketers about where they put their spend. Even previous algorithmic concerns like TikTok’s heating feature, where the company’s staff can boost videos to get them onto more feeds, did not alter advertisers’ spend on the platform.
“TikTok is a wildcard this year because of these [regulatory] concerns. Still, it’s the fastest-growing social network in terms of ad spend,” said Enberg.
Yandex scrapes Google and other SEO learnings from the source code leak
“Fragments” of Yandex’s codebase leaked online last week. Much like Google, Yandex is a platform with many aspects such as email, maps, a taxi service, etc. The code leak featured chunks of all of it.
According to the documentation therein, Yandex’s codebase was folded into one large repository called Arcadia in 2013. The leaked codebase is a subset of all projects in Arcadia and we find several components in it related to the search engine in the “Kernel,” “Library,” “Robot,” “Search,” and “ExtSearch” archives.
The move is wholly unprecedented. Not since the AOL search query data of 2006 has something so material related to a web search engine entered the public domain.
Although we are missing the data and many files that are referenced, this is the first instance of a tangible look at how a modern search engine works at the code level.
Personally, I can’t get over how fantastic the timing is to be able to actually see the code as I finish my book “The Science of SEO” where I’m talking about Information Retrieval, how modern search engines actually work, and how to build a simple one yourself.
In any event, I’ve been parsing through the code since last Thursday and any engineer will tell you that is not enough time to understand how everything works. So, I suspect there will be several more posts as I keep tinkering.
Also, shout out to Ryan Jones for digging in and sharing some key findings with me over IM.
OK, let’s get busy!
It’s not Google’s code, so why do we care?
Some believe that reviewing this codebase is a distraction and that there is nothing that will impact how they make business decisions. I find that curious considering these are people from the same SEO community that used the CTR model from the 2006 AOL data as the industry standard for modeling across any search engine for many years to follow.
That said, Yandex is not Google. Yet the two are state-of-the-art web search engines that have continued to stay at the cutting edge of technology.
Software engineers from both companies go to the same conferences (SIGIR, ECIR, etc) and share findings and innovations in Information Retrieval, Natural Language Processing/Understanding, and Machine Learning. Yandex also has a presence in Palo Alto and Google previously had a presence in Moscow.
A quick LinkedIn search uncovers a few hundred engineers that have worked at both companies, although we don’t know how many of them have actually worked on Search at both companies.
In a more direct overlap, Yandex also makes usage of Google’s open source technologies that have been critical to innovations in Search like TensorFlow, BERT, MapReduce, and, to a much lesser extent, Protocol Buffers.
So, while Yandex is certainly not Google, it’s also not some random research project that we’re talking about here. There is a lot we can learn about how a modern search engine is built from reviewing this codebase.
At the very least, we can disabuse ourselves of some obsolete notions that still permeate SEO tools like text-to-code ratios and W3C compliance or the general belief that Google’s 200 signals are simply 200 individual on and off-page features rather than classes of composite factors that potentially use thousands of individual measures.
Some context on Yandex’s architecture
Without context or the ability to successfully compile, run, and step through it, source code is very difficult to make sense of.
Typically, new engineers get documentation, walk-throughs, and engage in pair programming to get onboarded to an existing codebase. And, there is some limited onboarding documentation related to setting up the build process in the docs archive. However, Yandex’s code also references internal wikis throughout, but those have not leaked and the commenting in the code is also quite sparse.
Luckily, Yandex does give some insights into its architecture in its public documentation. There are also a couple of patents they’ve published in the US that help shed a bit of light. Namely:
As I’ve been researching Google for my book, I’ve developed a much deeper understanding of the structure of its ranking systems through various whitepapers, patents, and talks from engineers couched against my SEO experience. I’ve also spent a lot of time sharpening my grasp of general Information Retrieval best practices for web search engines. It comes as no surprise that there are indeed some best practices and similarities at play with Yandex.
Yandex’s documentation discusses a dual-distributed crawler system. One for real-time crawling called the “Orange Crawler” and another for general crawling.
Historically, Google is said to have had an index stratified into three buckets, one for housing real-time crawl, one for regularly crawled and one for rarely crawled. This approach is considered a best practice in IR.
Yandex and Google differ in this respect, but the general idea of segmented crawling driven by an understanding of update frequency holds.
One thing worth calling out is that Yandex has no separate rendering system for JavaScript. They say this in their documentation and, although they have Webdriver-based system for visual regression testing called Gemini, they limit themselves to text-based crawl.
The documentation also discusses a sharded database structure that breaks pages down into an inverted index and a document server.
Just like most other web search engines the indexing process builds a dictionary, caches pages, and then places data into the inverted index such that bigrams and trigams and their placement in the document is represented.
This differs from Google in that they moved to phrase-based indexing, meaning n-grams that can be much longer than trigrams a long time ago.
However, the Yandex system uses BERT in its pipeline as well, so at some point documents and queries are converted to embeddings and nearest neighbor search techniques are employed for ranking.
The ranking process is where things begin to get more interesting.
Yandex has a layer called Metasearch where cached popular search results are served after they process the query. If the results are not found there, then the search query is sent to a series of thousands of different machines in the Basic Search layer simultaneously. Each builds a posting list of relevant documents then returns it to MatrixNet, Yandex’s neural network application for re-ranking, to build the SERP.
Based on videos wherein Google engineers have talked about Search’s infrastructure, that ranking process is quite similar to Google Search. They talk about Google’s tech being in shared environments where various applications are on every machine and jobs are distributed across those machines based on the availability of computing power.
One of the use cases is exactly this, the distribution of queries to an assortment of machines to process the relevant index shards quickly. Computing the posting lists is the first place that we need to consider the ranking factors.
There are 17,854 ranking factors in the codebase
On the Friday following the leak, the inimitable Martin MacDonald eagerly shared a file from the codebase called web_factors_info/factors_gen.in. The file comes from the “Kernel” archive in the codebase leak and features 1,922 ranking factors.
Naturally, the SEO community has run with that number and that file to eagerly spread news of the insights therein. Many folks have translated the descriptions and built tools or Google Sheets and ChatGPT to make sense of the data. All of which are great examples of the power of the community. However, the 1,922 represents just one of many sets of ranking factors in the codebase.
A deeper dive into the codebase reveals that there are numerous ranking factor files for different subsets of Yandex’s query processing and ranking systems.
Combing through those, we find that there are actually 17,854 ranking factors in total. Included in those ranking factors are a variety of metrics related to:
Clicks.
Dwell time.
Leveraging Yandex’s Google Analytics equivalent, Metrika.
There is also a series of Jupyter notebooks that have an additional 2,000 factors outside of those in the core code. Presumably, these Jupyter notebooks represent tests where engineers are considering additional factors to add to the codebase. Again, you can review all of these features with metadata that we collected from across the codebase at this link.
Yandex’s documentation further clarifies that they have three classes of ranking factors: Static, Dynamic, and those related specifically to the user’s search and how it was performed. In their own words:
In the codebase these are indicated in the rank factors files with the tags TG_STATIC and TG_DYNAMIC. The search related factors have multiple tags such as TG_QUERY_ONLY, TG_QUERY, TG_USER_SEARCH, and TG_USER_SEARCH_ONLY.
While we have uncovered a potential 18k ranking factors to choose from, the documentation related to MatrixNet indicates that scoring is built from tens of thousands of factors and customized based on the search query.
This indicates that the ranking environment is highly dynamic, similar to that of Google environment. According to Google’s “Framework for evaluating scoring functions” patent, they have long had something similar where multiple functions are run and the best set of results are returned.
Finally, considering that the documentation references tens of thousands of ranking factors, we should also keep in mind that there are many other files referenced in the code that are missing from the archive. So, there is likely more going on that we are unable to see. This is further illustrated by reviewing the images in the onboarding documentation which shows other directories that are not present in the archive.
For instance, I suspect there is more related to the DSSM in the /semantic-search/ directory.
The initial weighting of ranking factors
I first operated under the assumption that the codebase didn’t have any weights for the ranking factors. Then I was shocked to see that the nav_linear.h file in the /search/relevance/ directory features the initial coefficients (or weights) associated with ranking factors on full display.
This section of the code highlights 257 of the 17,000+ ranking factors we’ve identified. (Hat tip to Ryan Jones for pulling these and lining them up with the ranking factor descriptions.)
For clarity, when you think of a search engine algorithm, you’re probably thinking of a long and complex mathematical equation by which every page is scored based on a series of factors. While that is an oversimplification, the following screenshot is an excerpt of such an equation. The coefficients represent how important each factor is and the resulting computed score is what would be used to score selecter pages for relevance.
These values being hard-coded suggests that this is certainly not the only place that ranking happens. Instead, this function is most likely where the initial relevance scoring is done to generate a series of posting lists for each shard being considered for ranking. In the first patent listed above, they talk about this as a concept of query-independent relevance (QIR) which then limits documents prior to reviewing them for query-specific relevance (QSR).
The resulting posting lists are then handed off to MatrixNet with query features to compare against. So while we don’t know the specifics of the downstream operations (yet), these weights are still valuable to understand because they tell you the requirements for a page to be eligible for the consideration set.
However, that brings up the next question: what do we know about MatrixNet?
There is neural ranking code in the Kernel archive and there are numerous references to MatrixNet and “mxnet” as well as many references to Deep Structured Semantic Models (DSSM) throughout the codebase.
The description of one of the FI_MATRIXNET ranking factor indicates that MatrixNet is applied to all factors.
Description: “MatrixNet is applied to all factors – the formula”
}
There’s also a bunch of binary files that may be the pre-trained models themselves, but it’s going to take me more time to unravel those aspects of the code.
What is immediately clear is that there are multiple levels to ranking (L1, L2, L3) and there is an assortment of ranking models that can be selected at each level.
The selecting_rankings_model.cpp file suggests that different ranking models may be considered at each layer throughout the process. This is basically how neural networks work. Each level is an aspect that completes operations and their combined computations yield the re-ranked list of documents that ultimately appears as a SERP. I’ll follow up with a deep dive on MatrixNet when I have more time. For those that need a sneak peek, check out the Search result ranker patent.
For now, let’s take a look at some interesting ranking factors.
Top 5 negatively weighted initial ranking factors
The following is a list of the highest negatively weighted initial ranking factors with their weights and a brief explanation based on their descriptions translated from Russian.
FI_ADV: -0.2509284637 -This factor determines that there is advertising of any kind on the page and issues the heaviest weighted penalty for a single ranking factor.
FI_DATER_AGE: -0.2074373667 – This factor is the difference between the current date and the date of the document determined by a dater function. The value is 1 if the document date is the same as today, 0 if the document is 10 years or older, or if the date is not defined. This indicates that Yandex has a preference for older content.
FI_QURL_STAT_POWER: -0.1943768768 – This factor is the number of URL impressions as it relates to the query. It seems as though they want to demote a URL that appears in many searches to promote diversity of results.
FI_COMM_LINKS_SEO_HOSTS: -0.1809636391 – This factor is the percentage of inbound links with “commercial” anchor text. The factor reverts to 0.1 if the proportion of such links is more than 50%, otherwise, it’s set to 0.
FI_GEO_CITY_URL_REGION_COUNTRY: -0.168645758 – This factor is the geographical coincidence of the document and the country that the user searched from. This one doesn’t quite make sense if 1 means that the document and the country match.
In summary, these factors indicate that, for the best score, you should:
Avoid ads.
Update older content rather than make new pages.
Make sure most of your links have branded anchor text.
Everything else in this list is beyond your control.
Top 5 positively weighted initial ranking factors
To follow up, here’s a list of the highest weighted positive ranking factors.
FI_URL_DOMAIN_FRACTION: +0.5640952971 – This factor is a strange masking overlap of the query versus the domain of the URL. The example given is Chelyabinsk lottery which abbreviated as chelloto. To compute this value, Yandex find three-letters that are covered (che, hel, lot, olo), see what proportion of all the three-letter combinations are in the domain name.
FI_QUERY_DOWNER_CLICKS_COMBO: +0.3690780393 – The description of this factor is that is “cleverly combined of FRC and pseudo-CTR.” There is no immediate indication of what FRC is.
FI_MAX_WORD_HOST_CLICKS: +0.3451158835 – This factor is the clickability of the most important word in the domain. For example, for all queries in which there is the word “wikipedia” click on wikipedia pages.
FI_MAX_WORD_HOST_YABAR: +0.3154394573 – The factor description says “the most characteristic query word corresponding to the site, according to the bar.” I’m assuming this means the keyword most searched for in Yandex Toolbar associated to the site.
FI_IS_COM: +0.2762504972 – The factor is that the domain is a .COM.
In other words:
Play word games with your domain.
Make sure it’s a dot com.
Encourage people to search for your target keywords in the Yandex Bar.
Keep driving clicks.
There are plenty of unexpected initial ranking factors
What’s more interesting in the initial weighted ranking factors are the unexpected ones. The following is a list of seventeen factors that stood out.
FI_PAGE_RANK: +0.1828678331 – PageRank is the 17th highest weighted factor in Yandex. They previously removed links from their ranking system entirely, so it’s not too shocking how low it is on the list.
FI_SPAM_KARMA: +0.00842682963 – The Spam karma is named after “antispammers” and is the likelihood that the host is spam; based on Whois information
FI_SUBQUERY_THEME_MATCH_A: +0.1786465163 – How closely the query and the document match thematically. This is the 19th highest weighted factor.
FI_REG_HOST_RANK: +0.1567124399 – Yandex has a host (or domain) ranking factor.
FI_URL_LINK_PERCENT: +0.08940421124 – Ratio of links whose anchor text is a URL (rather than text) to the total number of links.
FI_PAGE_RANK_UKR: +0.08712279101 – There is a specific Ukranian PageRank
FI_IS_NOT_RU: +0.08128946612 – It’s a positive thing if the domain is not a .RU. Apparently, the Russian search engine doesn’t trust Russian sites.
FI_YABAR_HOST_AVG_TIME2: +0.07417219313 – This is the average dwell time as reported by YandexBar
FI_LERF_LR_LOG_RELEV: +0.06059448504 – This is link relevance based on the quality of each link
FI_NUM_SLASHES: +0.05057609417 – The number of slashes in the URL is a ranking factor.
FI_ADV_PRONOUNS_PORTION: -0.001250755075 – The proportion of pronoun nouns on the page.
FI_TEXT_HEAD_SYN: -0.01291908335 – The presence of [query] words in the header, taking into account synonyms
FI_PERCENT_FREQ_WORDS: -0.02021022114 – The percentage of the number of words, that are the 200 most frequent words of the language, from the number of all words of the text.
FI_YANDEX_ADV: -0.09426121965 – Getting more specific with the distaste towards ads, Yandex penalizes pages with Yandex ads.
FI_AURA_DOC_LOG_SHARED: -0.09768630485 – The logarithm of the number of shingles (areas of text) in the document that are not unique.
FI_AURA_DOC_LOG_AUTHOR: -0.09727752961 – The logarithm of the number of shingles on which this owner of the document is recognized as the author.
FI_CLASSIF_IS_SHOP: -0.1339319854 – Apparently, Yandex is going to give you less love if your page is a store.
The primary takeaway from reviewing these odd rankings factors and the array of those available across the Yandex codebase is that there are many things that could be a ranking factor.
I suspect that Google’s reported “200 signals” are actually 200 classes of signal where each signal is a composite built of many other components. In much the same way that Google Analytics has dimensions with many metrics associated, Google Search likely has classes of ranking signals composed of many features.
Yandex scrapes Google, Bing, YouTube and TikTok
The codebase also reveals that Yandex has many parsers for other websites and their respective services. To Westerners, the most notable of those are the ones I’ve listed in the heading above. Additionally, Yandex has parsers for a variety of services that I was unfamiliar with as well as those for its own services.
What is immediately evident, is that the parsers are feature complete. Every meaningful component of the Google SERP is extracted. In fact, anyone that might be considering scraping any of these services might do well to review this code.
There is other code that indicates Yandex is using some Google data as part of the DSSM calculations, but the 83 Google named ranking factors themselves make it clear that Yandex has leaned on the Google’s results pretty heavily.
Yandex has anti-SEO upper bounds for some ranking factors
315 ranking factors have thresholds at which any computed value beyond that indicates to the system that that feature of the page is over-optimized. 39 of these ranking factors are part of the initially weighted factors that may keep a page from being included in the initial postings list. You can find these in the spreadsheet I’ve linked to above by filtering for the Rank Coefficient and the Anti-SEO column.
It’s not far-fetched conceptually to expect that all modern search engines set thresholds on certain factors that SEOs have historically abused such as anchor text, CTR, or keyword stuffing. For instance, Bing was said to leverage the abusive usage of the meta keywords as a negative factor.
Yandex boosts “Vital Hosts”
Yandex has a series of boosting mechanisms throughout its codebase. These are artificial improvements to certain documents to ensure they score higher when being considered for ranking.
Below is a comment from the “boosting wizard” which suggests that smaller files benefit best from the boosting algorithm.
There are several types of boosts; I’ve seen one boost related to links and I’ve also seen a series of “HandJobBoosts” which I can only assume is a weird translation of “manual” changes.
One of these boosts I found particularly interesting is related to “Vital Hosts.” Where a vital host can be any site specified. Specifically mentioned in the variables is NEWS_AGENCY_RATING which leads me to believe that Yandex gives a boost that biases its results to certain news organizations.
Without getting into geopolitics, this is very different from Google in that they have been adamant about not introducing biases like this into their ranking systems.
The structure of the document server
The codebase reveals how documents are stored in Yandex’s document server. This is helpful in understanding that a search engine does not simply make a copy of the page and save it to its cache, it’s capturing various features as metadata to then use in the downstream rankings process.
The screenshot below highlights a subset of those features that are particularly interesting. Other files with SQL queries suggest that the document server has closer to 200 columns including the DOM tree, sentence lengths, fetch time, a series of dates, and antispam score, redirect chain, and whether or not the document is translated. The most complete list I’ve come across is in /robot/rthub/yql/protos/web_page_item.proto.
What’s most interesting in the subset here is the number of simhashes that are employed. Simhashes are numeric representations of content and search engines use them for lightning fast comparison for the determination of duplicate content. There are various instances in the robot archive that indicate duplicate content is explicitly demoted.
Also, as part of the indexing process, the codebase features TF-IDF, BM25, and BERT in its text processing pipeline. It’s not clear why all of these mechanisms exist in the code because there is some redundancy in using them all.
Link factors and prioritization
The codebase also reveals a lot of information about link factors and how links are prioritized.
Yandex’s link spam calculator has 89 factors that it looks at. Anything marked as SF_RESERVED is deprecated. Where provided, you can find the descriptions of these factors in the Google Sheet linked above.
Notably, Yandex has a host rank and some scores that appear to live on long term after a site or page develops a reputation for spam.
Another thing Yandex does is review copy across a domain and determine if there is duplicate content with those links. This can be sitewide link placements, links on duplicate pages, or simply links with the same anchor text coming from the same site.
This illustrates how trivial it is to discount multiple links from the same source and clarifies how important it is to target more unique links from more diverse sources.
What can we apply from Yandex to what we know about Google?
Naturally, this is still the question on everyone’s mind. While there are certainly many analogs between Yandex and Google, truthfully, only a Google Software Engineer working on Search could definitively answer that question.
Yet, that is the wrong question.
Really, this code should help us expand our thinking about modern search. Much of the collective understanding of search is built from what the SEO community learned in the early 2000s through testing and from the mouths of search engineers when search was far less opaque. That unfortunately has not kept up with the rapid pace of innovation.
Insights from the many features and factors of the Yandex leak should yield more hypotheses of things to test and consider for ranking in Google. They should also introduce more things that can be parsed and measured by SEO crawling, link analysis, and ranking tools.
For instance, a measure of the cosine similarity between queries and documents using BERT embeddings could be valuable to understand versus competitor pages since it’s something that modern search engines are themselves doing.
Much in the way the AOL Search logs moved us from guessing the distribution of clicks on SERP, the Yandex codebase moves us away from the abstract to the concrete and our “it depends” statements can be better qualified.
To that end, this codebase is a gift that will keep on giving. It’s only been a weekend and we’ve already gleaned some very compelling insights from this code.
I anticipate some ambitious SEO engineers with far more time on their hands will keep digging and maybe even fill in enough of what’s missing to compile this thing and get it working. I also believe engineers at the different search engines are also going through and parsing out innovations that they can learn from and add to their systems.
Simultaneously, Google lawyers are probably drafting aggressive cease and desist letters related to all the scraping.
I’m eager to see the evolution of our space that’s driven by the curious people who will maximize this opportunity.
But, hey, if getting insights from actual code is not valuable to you, you’re welcome to go back to doing something more important like arguing about subdomains versus subdirectories.
Opinions expressed in this article are those of the guest author and not necessarily Search Engine Land. Staff authors are listed here.
@media screen and (min-width: 800px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:770px; min-height:260px; }
}
@media screen and (min-width: 1279px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:800px!important; min-height:440px!important; }
}
About the author
An artist and a technologist all rolled into one, Michael King is the CEO and founder of digital marketing agency, iPullRank focused on technical SEO, content strategy and machine learning. King consults with enterprise and mid-market companies all over the world, including brands like SAP, American Express, Nordstrom, SanDisk, General Mills, and FTD. King’s background in Computer Science and as an independent hip-hop musician sets him up for deep technical and creative solutions for modern marketing problems. Check out his book “The Science of SEO” coming in Q2 of 2023 from Wiley Publishing.
Baidu working on AI chatbot service that will be added to search
First Microsoft Bing. Then Google. Now Baidu is reportedly planning on bringing ChatGPT-style AI to its search results.
Why we care. All the major search engines are seemingly in an arms race to add AI chat to search. Once search engines eventually add the chat features to search, it could have major implications for publishers (websites could see their traffic and visibility impacted, depending on how the AI chat is deployed within the search results) and searchers (will the information be accurate and reliable?). There are a lot of unknown unknowns here, which means search marketers should be watching all these developments.
A standalone app first, then search. Baidu is expected to launch its AI chatbot first as a standalone app (similar to ChatGPT). It would then be gradually merged into Baidu search by March, according to reports.
Baidu is reportedly using its deep learning model called ERNIE (which Baidu described as “a “pre-training language model with 260 billion parameters”) as the chatbot’s foundation and “training it on both Chinese- and English-language sources inside and outside China’s firewall,” according to the Wall Street Journal. Baidu also will limit the outputs of its chatbot to comply with China’s censorship rules.
Dig deeper. There’s more coverage of the news on Techmeme.
@media screen and (min-width: 800px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:770px; min-height:260px; }
}
@media screen and (min-width: 1279px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:800px!important; min-height:440px!important; }
}
About the author
Danny Goodwin is Managing Editor of Search Engine Land & SMX. In addition to writing daily about SEO, PPC, and more for Search Engine Land, Goodwin also manages Search Engine Land’s roster of subject-matter experts. He also helps program our conference series, SMX – Search Marketing Expo. Prior to joining Search Engine Land, Goodwin was Executive Editor at Search Engine Journal, where he led editorial initiatives for the brand. He also was an editor at Search Engine Watch. He has spoken at many major search conferences and virtual events, and has been sourced for his expertise by a wide range of publications and podcasts.
Yahoo Search seems like it will be making a comeback in the future. Yahoo has been dropping hints over the past couple of weeks related to this return and is also hiring a Principal Product Manager for the Yahoo Search platform to help lead these initiatives.
The job posting. Yahoo posted a job listing for a “Principal Product Manager, Yahoo Search” a few weeks ago. The job posting, in part, reads, “We’re looking for a Product Manager for Search at Yahoo. We are looking for folks that are interested in pushing beyond the status quo to change the way folks interact and use search.”
“As a Product Manager for Search, you will help develop our search strategy and roadmap and lead its execution. The ideal candidate will leverage strong organizational skills and deep subject matter expertise to partner with design, science, engineering, and other key cross-functional teams. You will determine what we prioritize for our customers in our search experiences and bring the vision to life. You will also lead the effort to discover and amplify content from across the vast Yahoo ecosystem to create new and innovative search experiences across surfaces and for our Search App. The role is also responsible for identifying and documenting product and business requirements and taking them from concept to production, while working with a broad set of stakeholders that include marketing, sales, legal, editorial, design, UXR, and other teams,” it continues to read.
Twitter hints. Yahoo has reactivated its Twitter account for Yahoo Search, posting teasers throughout the past couple of weeks. Here are some of those:
Just popping in to remind everyone that we did search before it was cool.
Yahoo executives. Brian Provost, SVP & GM, Yahoo, posted on LinkedIn about this job listing and wrote, “There’s going to be so much innovation in Search in the coming years and there aren’t many places where you can immediately have an impact this big. Would love to hear from you if you have a passion for Search and building product experiences.”
Karen Chin, Sr. Director of Product Management, Yahoo, also posted on LinkedIn, saying, “Looking to drive meaningful and innovative experiences for millions of users? We are looking for a seasoned Search Product Manager to take search into the next phase! Share and join us.”
Jim Lanzone, Chief Executive Officer at Yahoo, took the helm of Yahoo in September 2021. Jim has a lot of deep roots in search. He worked at Ask.com for seven years, starting in 2001 as an SVP, Product Management, then in 2004 as the SVP and GM of Ask Jeeves and then taking over as CEO in 2006. After Ask.com, he became the President and CEO of CBS Interactive, then the CEO at Tinder and now at Yahoo as their CEO. It will be exciting to see what Yahoo Search does under Jim’s leadership. He is a creative mind that produced a lot of search innovation at Ask.
Why we care. Personally, I cannot wait to see what Jim and his team come up with for Yahoo Search. I am excited to see what new ideas, interfaces, and concepts the team brings to Yahoo Search. Yahoo was a pretty big player in search in the early days, then the company continued to decline and even Google veteran Marissa Mayer could not save the company.
But now Yahoo has a blank slate, and it will be very exciting to see if Yahoo can compete again.
@media screen and (min-width: 800px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:770px; min-height:260px; }
}
@media screen and (min-width: 1279px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:800px!important; min-height:440px!important; }
}
About the author
Barry Schwartz a Contributing Editor to Search Engine Land and a member of the programming team for SMX events. He owns RustyBrick, a NY based web consulting firm. He also runs Search Engine Roundtable, a popular search blog on very advanced SEM topics. Barry can be followed on Twitter here.
A former employee allegedly leaked a Yandex source code repository, part of which contained more than 1,900 factors used by the search engines for ranking websites in search results.
Why we care. This leak has revealed 1,922 ranking factors Yandex used in its search algorithm, at least as of July 2022. Perhaps Martin MacDonald put it best on Twitter today: “The Yandex hack is probably the most interesting thing to have happened in SEO in years.”
Yandex is not Google. If you plan to read the full list of Yandex ranking factors, remember that Yandex is not Google. If you see a ranking factor listed by Yandex, that doesn’t mean Google gives that signal that same amount of weight. In fact, Google may not use all of the 1,922 factors listed. In fact, many of the factors in this leak are deprecated or unused.
That said, a lot of these ranking factors may be quite similar to signals Google uses for search. So reviewing this document may provide some useful insights to better help you understand how search engines, such as Google, work from a technological standpoint.
The bigger picture. The code appeared as a Torrent on a popular hacking forum, as reported by Bleeping Computer:
…the leaker posted a magnet link that they claim are ‘Yandex git sources’ consisting of 44.7 GB of files stolen from the company in July 2022. These code repositories allegedly contain all of the company’s source code besides anti-spam rules.
Yandex calls it a leak. Because the code appeared on a popular hacking forum, it was first thought that Yandex was hacked. Yandex has denied this, and provided the following statement:
“Yandex was not hacked. Our security service found code fragments from an internal repository in the public domain, but the content differs from the current version of the repository used in Yandex services.
A repository is a tool for storing and working with code. Code is used in this way internally by most companies.
Repositories are needed to work with code and are not intended for the storage of personal user data. We are conducting an internal investigation into the reasons for the release of source code fragments to the public, but we do not see any threat to user data or platform performance.”
Dig deeper. You can find more coverage of the leak on Techmeme.
Yandex ranking factors list. MacDonald shared the full list of 1,922 factors here on Web Marketing School. I highly recommend downloading it, as I fully expect Yandex will try to scrub this information from the internet. (Editor’s note: In an earlier version of this article, we had linked to a translated version on Dropbox, but that link quickly went away.)
Early analysis of ranking factors. Alex Buraks created two Twitter threads – first thread, second thread – analyzing the various ranking factors. There’s another interesting Twitter thread here from Michael King.
Many of Yandex’s ranking factors are what you’d expect to see:
PageRank and many link-related factors (e.g., age, relevancy, etc.).
Text relevancy.
Content age and freshness.
End-user behavior signals.
Host reliability.
Some sites get preference (e.g., Wikipedia).
Some of the ranking factors SEOs are finding surprising: number of unique visitors, percent of organic traffic and average domain ranking across queries.
And as Taylor pointed out, 244 of the ranking factors were categorized as unused and 988 as deprecated, “meaning that 64% of the document is either not actively used or has been superseded – so it’s more like ~690 potential ranking factors, and a lot of them contain thin descriptions.”
@media screen and (min-width: 800px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:770px; min-height:260px; }
}
@media screen and (min-width: 1279px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:800px!important; min-height:440px!important; }
}
About the author
Danny Goodwin is Managing Editor of Search Engine Land & SMX. In addition to writing daily about SEO, PPC, and more for Search Engine Land, Goodwin also manages Search Engine Land’s roster of subject-matter experts. He also helps program our conference series, SMX – Search Marketing Expo. Prior to joining Search Engine Land, Goodwin was Executive Editor at Search Engine Journal, where he led editorial initiatives for the brand. He also was an editor at Search Engine Watch. He has spoken at many major search conferences and virtual events, and has been sourced for his expertise by a wide range of publications and podcasts.
Convergent TV Summit returns March 21-22. Hear timely insights from TV industry experts virtually or in person in NYC. Register now to secure your early bird pass.
TikTok now allows users to receive direct messages from every other user on the platform, even if they aren’t following them.
In order to receive direct messages from everyone, users will need to change a setting in the application’s Privacy menu.
Our guide will show you how to receive direct messages from everyone in the TikTok mobile app.
Note: These screenshots were captured in the TikTok app on iOS.
Step 1: On your TikTok profile, tap the three horizontal lines in the top-right corner of the screen.
Step 2: Tap “Settings and privacy.”
Step 3: Tap “Privacy.”
Step 4: Under the “Interactions” section, tap “Direct messages.”
Step 5: Tap “Direct messages” on the left side of the screen.
Step 6: Tap “Everyone” on the window that appears at the bottom of the screen.
Step 7: Tap the “x” icon in the top-right corner of the “Direct messages” window to close the window and continue using the TikTok app.
Convergent TV Summit returns March 21-22. Hear timely insights from TV industry experts virtually or in person in NYC. Register now to secure your early bird pass.
TikTok allows users to receive push notifications when certain actions take place in the photo- and video-sharing application. For instance, users can receive a notification when another user sends them a direct message.
The TikTok app allows users to change their push notifications settings so they’ll only receive the notifications they’re interested in.
Our guide will show you how to change your push notifications settings in the TikTok mobile app.
Note: These screenshots were captured in the TikTok app on iOS.
Step 1: On your TikTok profile, tap the three horizontal lines in the top-right corner of the screen.
Step 2: Tap “Settings and privacy” at the bottom of the screen.
Step 3: Under the “Content & Display” section, tap “Notifications.”
Step 4: You’ll see a list of the different notifications you can receive. If the toggle to the right of a notification setting is blue, the notification setting is turned “on.” If the toggle is gray, the notification setting is turned “off.” Tap the toggle to the right of each notification option to turn the notification on or off, depending on your preference.
This guide was first published in March 2019 and was updated in January 2023.
SEO professionals can access many free and paid platforms, tools, and software. But if you and your competitors are all using the same tools, data, and approaches – how can you set yourself apart?
At SMX Next, I shared the SEO tools that will make up my toolkit in 2023. If you missed the session, read on as I share some highlights.
Before we dive in, it’s worth noting that true to the fast pace of SEO, there’s already been a change in how you might think of one of the tools – but we’ll get to that shortly.
So what will we be covering in this article?
Why do we use SEO tools?
The usual suspects.
The unusual suspects.
Fun with AI.
Let’s get started.
As with virtually every decision we make, when it comes to tools, it’s good to think about the ever-present question – why? Why do we use tools in the first place?
We generally use SEO tools for one of the following tasks:
To automate monotonous tasks (e.g., rank checking).
To distill large amounts of data into usable information (e.g., Google Analytics).
To access information not functionally available to us otherwise (e.g., backlink analysis).
To combine data sources and points (e.g., domain metrics on backlinks).
Essentially, we use tools to save time. When you think of most of the tools you use or want to use, they generally speed up the collection of information or present it in a way that’s easier to draw conclusions.
Take a moment and think of one of the tools you use and one of the tasks you use it for. Guaranteed it will fit one or more of the purposes above.
But, of course, there are hundreds of tools to choose from. I’m not going to pretend that the list below is fully exhaustive.
I’ve only included tools I use myself and only those that would apply to most people. That said, they’re tools that have proven themselves invaluable to my regular routine.
So what are they?
The usual suspects
You can probably guess what my list of usual suspects is now. But let’s cover them anyway.
Semrush
Why does Semrush make this list?
Semrush is a solid all-in-one toolset covering multiple areas of SEO, SEM and SMM and generally does them well.
Some of the key tools within the suite I use regularly are:
Audit: There are definitely more sophisticated technical audit tools out there, but the Semrush tool runs regularly and gives you a quick-and-easy place to monitor for spikes and drops in errors, etc.
Rank checking: Their rank-checking tool is excellent. You can select the location your rankings are checked from (or multiple) and monitor them daily. For those with large numbers of terms, they’re also presented in an easily digested and filterable way.
Social poster: Semrush has a solid social monitoring and posting tool. You can keep track of your progress against competitors and monitor RSS feeds to quickly and easily get post ideas. It’s not as robust a social tool as some dedicated ones are, but it’s solid and good enough for many use cases.
You’re probably familiar with much of what it does and if not, I’d suggest a trial.
Semrush is generally pretty affordable for the variety of purposes it serves (though it can get pricey if you need to add users, white label reports, or use add-ons). And since it’s generally well known, knowing how to use it is a marketable skill.
Screaming Frog
Hands down, this is the best value tool on the market.
Screaming Frog, for the three of you who might not know, is a crawler.
Give it a website or a list of URLs and configure how you want it to crawl (depth, user agents, paths, etc.), info on the data you want to collect, plus a bit of time, and it’ll return an audit of the site with various visualizations and filters.
There’s a free version that’s good for up to 500 URLs, though it has limited customizations. For a paid version, you’ll have to pony up a “whopping” $210/year. As I said, best value tool on the market.
It’s good for:
Customized crawling.
Looking for specific text or HTML on pages.
Easily export issues to send to devs.
Creates XML sitemaps.
Great visualizations.
And I don’t have to dive into the remainder of my usual suspects as you hopefully use them already:
All I’ll say about Bing Webmaster Tools is this:
It’s like Search Console, but with far more details and information, making it the unsung hero of SEO tools.
The unusual suspects
Technical SEO tools
Jet Octopus is another technical SEO suite. (And no, I don’t know where these companies get their names.)
The interface includes screens like:
It’s similar to Semrush but with different filtering options and of course, a different crawler. As you can see above, Jet Octopus lets you easily group issues into the sections of the site you’ll find them in, like a combination of Semrush and Bing Webmaster Tools.
I also find it gives me a different way of looking at a site structure, though I wouldn’t give up Semrush for it, making it one to add to the mix if you have the budget and need to make sure you have an easy-to-use different way of looking at things.
And some additional unusual suspects for technical SEO:
Merkle: Various free SEO tools covering everything from schema to prerendering.
Structured Data Testing Tool: For testing your schema.
Mobile Moxie: Free and paid tools for testing your mobile SEO.
Uptime Robot: For making sure your site(s) are up.
Content tools
But all the technical SEO in the world won’t get your ranking without great content. So let’s look at some unusual suspects on the content side:
A huge favorite content-related SEO tool of mine is the shockingly inexpensive Infranodus (though it arguably isn’t a content tool, it’s what I use it most for).
Within Infranodus, you’ll find an array of tools for a whopping €9/month (about the same in USD).
My most commonly used tools help me dig into the concepts included in the top Google results, and how the concepts connect to various pages within a list.
In short, you can enter a query and it will produce an interactive mapping of how the terms in the results connect to each other, which I find helps me not only better understand how Google might see a concept, but a user as well.
Here it is in action:
And of course:
AlsoAsked: A good visualization of questions people as related to topics you’re researching.
Answer The Public: A good visualization of how the questions related to a phrase group, by question intent.
Link-related tools
Of course, I have tools I use for keeping an eye on links and competitors’ links.
My favorite tools in this category are:
Ahrefs
While technically Ahrefs is a suite of tools one might compare to Semrush, it’s in links that I find it really shines.
I find it catches new backlinks faster than other tools, and generally has a more robust database.
So when I’m keeping my eye on new links, looking for gaps in link profiles, or just researching competitors, this is my first (though not only) stop.
Majestic
I haven’t used Majestic in a few years, but I wanted to include it in the list as it’s a solid backlink tool worth your consideration.
At one point, I hit a threshold and had too many tools, so I made a “one in, one out” policy to keep it under control.
I had a lot of duplication in the link category so I had to get rid of some tools. But for those who don’t have this problem, Majestic is a solid option to consider (and might even be worth revisiting myself).
Search Console
I hope I don’t have to tell you why Search Console is an important tool for monitoring your backlinks, but here it is summed up in a Stable Diffusion-generated image.
Get the daily newsletter search marketers rely on.
Of course, we can’t forget about AI. There are all sorts of AI-driven tools and I’m not going to tell you which is best as I haven’t tested them all and am far from deciding which one(s) I’ll land on yet (though Jasper has taken an early lead).
That said, you can access some of the core tech for free!
While there’s been a lot of hype around ChatGPT, I still prefer accessing the technology (GPT-3) in the OpenAI playground. It’s the same technology with (IMO) better flexibility and interface.
Simply sign up for an OpenAI account, and then you can access the API to perform all sorts of NLP-related tasks, or just play around in their playground.
It can be used for:
Content outlines.
Writing ecommerce data en masse.
Powering chatbots (though there are many pre-boxed solutions for that as well).
Translation.
And so much more, including full content creation. (Use at your own risk!)
AI tools for image generation are also popular nowadays. I’ve used text-to-image generators for:
It lets you connect with your Google Analytics (only Universal Analytics for now, as most people don’t have over a year of data in GA4, which is required for decent forecasting).
You can even choose a specific segment of your analytics (organic, for example) and forecast the next few months to give you something like:
It’s nice to be able to know in advance what things are going to look like in the future.
@media screen and (min-width: 800px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:770px; min-height:260px; }
}
@media screen and (min-width: 1279px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:800px!important; min-height:440px!important; }
}
About the author
Dave Davies is the Lead SEO for the Machine Learning Operations company Weights & Biases. He got his start in SEO in the early 2000s and in 2004 co-founded Beanstalk Internet Marketing with his wife Mary, who still runs its day-to-day operations. He hosts a weekly podcast, speaks regularly at the industry’s leading conferences, and is proud to be a regular contributor right here on Search Engine Land.
Moving on from search engine optimization to search optimization
In the early 2000s, web search was mostly limited to search engines and online directories available at the time. But we have come a long way since then.
Optimizing for search remains important today. But users no longer solely rely on traditional search engines to look up information.
I believe it’s time for SEO professionals to think beyond Google and search engines and start looking at search holistically.
Beyond Google: The reality of today’s search landscape
As the competition online increases, the best way to do right by our clients is to ensure we’re working to improve their visibility, not just on search engines like Google, but also on social media platforms where their target audiences are searching.
Nowadays, a substantial share of web searches happens on social media platforms. More than social networking, these websites and apps are where many of today’s web users knowingly – or unknowingly – search via hashtags, trending topics and the like.
While we’ve gotten used to the term “SEO,” it’s about time we talk about “search optimization” instead of the limiting concept of “search engine optimization.” Search campaigns should focus on overall search presence rather than only giving importance to search engine results.
Yes, there is already a concept of social media optimization (SMO). But this generally refers to creating and sharing content on social media with the goal of making it viral.
What I am talking about is optimizing for searches done by social media users within social media platforms that have search functionality.
The mere fact that social media platforms have search features that are constantly enhanced is compelling proof. Brands must take their presence on social media search seriously – and as search marketers, our strategies should help set them up for long-term success.
Get the daily newsletter search marketers rely on.
Social search: What it is and why it matters
“Social search uses elements of user behavior, implicit and explicit, to improve the results of searches inside and outside enterprises. Such elements are typically stored as metadata, making social search a type of metadata mining. It also enables users to disambiguate results from their queries more effectively,” according to Gartner.
Showing up in social media search results is vital for brands. Here are 12 reasons why.
Social media platforms have allocated decent budgets to further improve their search functionality. Because of this, social search results will likely improve as well over time.
People are spending more time on social media than on search engines. Though Google search is still of great importance, the use of social search has increased alongside social media usage.
Gen Z users in the U.S. spend ~5 hours per week on Instagram.
Social search results have many filters like hashtags, services, companies, jobs, events, people, etc. Hence, the results are more refined and targeted.
Social platforms also focus on local business discoverability.
Social media search results show interaction and reflect word of mouth in the form of comments, likes, shares, follows and conversations.
Studies show that people trust what other people (clients, customers, and suppliers) have to say rather than what business owners display on their websites.
While search engine results typically surface optimized webpages, social search mainly focuses on relevance and engagement.
Since social media search results return content from posts and interactions of social media users, there is less spam than typical search engine results.
Quality content becomes more valuable when people talk about your products, services and brands on social media, and when these conversations and feedback appear in social media search.
Search engine results generally provide information. Social search results can be influential.
People of all ages get easily accustomed to using new trends in social media, but understanding search engines and adapting to advanced techniques on search engines can be more challenging.
To get a better idea, let’s explore the search features of popular social media platforms, such as:
Twitter
LinkedIn
Instagram
YouTube
Twitter
There are many ways to use the search function on Twitter. You can:
Find tweets from yourself, friends, local businesses, well-known entertainers, global political leaders and many more.
Search using search queries or use hashtags.
Follow ongoing conversations about the news, products, services, or any other personal interests.
Just like traditional search engine results, Twitter search results for users who are logged in differ from the ones who are not.
Advanced search is also available when you’re logged in, letting you get customized search results for specific date ranges, people, and more. This makes it easier to find specific tweets.
The platform also has a FAQs page about search results displayed on Twitter.
LinkedIn
LinkedIn is the world’s largest professional network on the web. It is a platform for anyone looking for varied career opportunities, including people from various professional backgrounds, such as small business owners, students, and job seekers.
LinkedIn members can use the platform to tap into a network of professionals, companies, and groups within and beyond their industry.
LinkedIn search has great filters to help narrow your search for more relevant results. You can use filters like:
People
Companies
Jobs
Posts
Groups
Events
Courses
Schools
Services
Additionally, the platform is widely used by job seekers and recruiters. HR professionals use it for headhunting and background checks on applicants.
LinkedIn is surely a social media platform, but with a focus on professional interactions, as well as knowledge-sharing.
From a search perspective, people use LinkedIn for professional development and seeking job opportunities. Hence, being found in the search results on LinkedIn is crucial.
Instagram
In 2021, there were 1.21 billion monthly active users on Meta’s Instagram, making up over 28% of the world’s internet users. By 2025, its user base is expected to grow to 1.44 billion, accounting for 31.2% of global internet users.
Hence, sharing innovative content regularly and using the correct hashtags can surely boost the chance of your profile showing up on Instagram search results.
Instagram content is usually in the form of images, videos, live conversations, reels, and stories. (And there’s a shopping feature, too.)
As a rapidly growing platform, you simply cannot ignore this social media platform – especially if you want to be in front of Gen Z users.
Instagram interactions are measured in likes, shares, and follows. If your client is in retail, fashion, food, baby products or grooming, then this is a place where their prospective buyers are searching for options.
YouTube
YouTube search prioritizes three main elements to provide the best search results:
Relevance
Engagement
Quality
Optimizing for YouTube search is like optimizing a website, but here the content is in video format. Thus, your video title and description have to be very relevant so that they can match the search query of the user.
On YouTube, prospective buyers can discover your brand. You can also create videos that answer frequently asked questions customers may have after buying your product or service. Hence, making it a great platform for sales and after-sales.
Optimizing for search beyond search engines
People are no longer limited to searching on search engines and online directories. Multiple tools are at their disposal to satisfy their need for information.
Thus, our search campaigns should not merely revolve around showing up on search engines like Google – but for every type of search, regardless of the platform.
In the past, we have seen how the search on directories became redundant. The future may see popular social media platforms having more sophisticated search capabilities – and people using them more than the actual search engines.
To future-proof our clients’ businesses and set them up for long-term success, start optimizing for search where their target audiences are – not just on Google but also on social media platforms.
Opinions expressed in this article are those of the guest author and not necessarily Search Engine Land. Staff authors are listed here.
@media screen and (min-width: 800px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:770px; min-height:260px; }
}
@media screen and (min-width: 1279px) { #div-gpt-ad-3191538-7 { display: flex !important; justify-content: center !important; align-items: center !important; min-width:800px!important; min-height:440px!important; }
}
About the author
Bharati Ahuja is the Founder of WebPro Technologies LLP. She is also an SEO Trainer and Speaker, Blog Writer, and Web Presence Consultant, who first started optimizing websites in 2000. Since then, her knowledge about SEO has evolved along with the evolution of search on the web.
Convergent TV Summit returns March 21-22. Hear timely insights from TV industry experts virtually or in person in NYC. Register now to secure your early bird pass.
The For You feed on TikTok is valuable real estate, and there are countless online posts featuring advice and best practices on how to stake a claim to part of that land, but the most foolproof method of doing so is apparently out of the control of brands and creators.
Emily Baker-White of Forbes spoke with six current and former employees of TikTok and parent company ByteDance and reviewed internal communications and documents confirming the existence of a practice referred to within the companies as “heating.”