Stars Are an Adversarial Metric: A Goal-Aware Credibility Model for Open Source Projects
Introduction
When an organization launches an open source project, the first measurement question is usually the same: how will we know whether it is succeeding, and how do we tie that success to an objective and a meaningful key result? The most common answer is GitHub stars. Stars appear in README badges, conference slides, investor decks, architecture reviews, procurement discussions, and countless awesome-* lists. They are easy to understand, easy to compare, and easy to report in a quarterly business review. A project with 50,000 stars looks successful.
This paper argues that stars are not merely a weak signal but an adversarial one. Once a metric influences discovery, hiring, funding, architecture decisions, or Developer Relations (DevRel) reporting, it creates an incentive for manipulation. Recent measurement research on suspected fake GitHub stars provides empirical grounding for this claim: He et al. describe stars as the most widely used GitHub popularity signal while showing that artificial inflation degrades their value as a decision-making input and can create security risk for GitHub users [1]. Stars can be bought, exchanged, incentivized, driven by hype, or accumulated because a project was fashionable for two weeks and then quietly abandoned. More often than not, they measure attention rather than trust, adoption, or long-term value. Forks fare little better: empirical studies of GitHub data show that most forks never produce contributions [12], so a fork may indicate intent to contribute, intent to inspect, or nothing at all.
If stars are not enough, what should we measure instead? And if we are setting a DevRel objective around an open source project, which metrics actually indicate that we are building something valuable?
Answering that question requires separating three constructs that everyday usage conflates. Quality is intrinsic: maintenance discipline, security posture, release hygiene, documentation. Adoption is extrinsic: dependents, downloads, integrations, contributors. Success is neither — it is the goal-relative attainment of whatever the project set out to do, and the two previous constructs combine into it differently for every project. A project can be a runaway success ahead of its quality maturity, as viral developer tools periodically demonstrate, and a polished, well-engineered repository can fully achieve its purpose — supporting a talk, documenting a pattern — while registering near zero on every adoption metric. No fixed scalar can measure success, because success has no fixed definition. What an external evaluator can measure is credibility: the decision-relevant question of whether a project can responsibly be depended upon. That is the construct this paper targets, and the model it proposes is goal-aware by design — its weight vector is precisely where an evaluator encodes how much quality and how much adoption their own definition of success requires. Credibility, on this account, is neither necessary for popularity-style success nor sufficient for goal attainment; it is the evidence layer underneath both.
This paper makes three contributions. First, it formalizes the argument that GitHub popularity metrics should be treated as adversarial signals, synthesizing recent measurement evidence [1], platform policy [2], and the classical Goodhart/Strathern analysis of metrics that become targets [17]. Second, it proposes the Open Source Credibility Coefficient (OSCC), a normalized, parameterized model of project credibility in which popularity signals are explicitly discounted by an estimated authentic-engagement fraction, health signals are aggregated with recency weighting, and structural risks (maintainer concentration, abandonment, unaddressed vulnerabilities) enter as bounded penalties. The model is presented as a structured review framework and a thinking tool, not as a production ranking algorithm; its weights are meant to be tuned to the evaluation context. Third, it derives an evidence hierarchy for open source metrics and shows how that hierarchy can be used to write DevRel objectives and key results that resist gaming and better reflect long-term value.
The remainder of the paper is organized as follows. Section 2 positions the work against existing community-health frameworks, composite scorers, and measurement studies. Section 3 develops the argument that stars are an adversarial metric. Section 4 presents the OSCC model. Section 5 operationalizes the suspicion factors for stars and forks. Section 6 presents the evidence hierarchy and its application to DevRel objectives. Section 7 illustrates the framework on four project archetypes. Section 8 discusses the reflexive limits of the model itself, and Section 9 states limitations and threats to validity. Section 10 concludes.
Related Work
This paper sits at the intersection of four research threads: the long-running effort to define and measure open source success and community health; empirical studies of GitHub popularity metrics and their semantics; composite scoring systems for open source projects; and the more recent literature on metric manipulation and adversarial signals.
Defining and measuring open source success
The question of what “success” means for an open source project predates GitHub. Crowston and Howison argued early on that classical information-systems success measures translate poorly to community-driven development, and proposed health indicators centered on process and community structure rather than output alone [7]. Jansen extended the discussion from individual projects to ecosystems, arguing that project health cannot be assessed in isolation from the network of projects, contributors, and consumers around it [8]. The most mature practical descendant of this line of work is the CHAOSS project under the Linux Foundation, which maintains a curated catalog of community-health metrics and metric models covering responsiveness, contributor diversity, organizational participation, and risk [9]. The framework proposed in this paper is deliberately compatible with that tradition: most of the time-dependent terms in the OSCC core correspond to metric families that CHAOSS already defines. What CHAOSS does not provide is a stance on how those metrics should be combined, weighted, decayed over time, or defended against manipulation, which is precisely the gap this paper targets.
The semantics of GitHub popularity metrics
A second thread examines what popularity metrics on social coding platforms actually mean. Borges and Valente surveyed hundreds of developers about why they star repositories and found that starring conflates appreciation, bookmarking, and signaling, with a substantial share of stars expressing intent rather than usage [10]. Earlier work by the same group studied the factors that drive repository popularity and showed that star growth is strongly shaped by promotion events and external visibility rather than by intrinsic project quality alone [11]. Kalliamvakou et al. documented broader perils of mining GitHub data, including the observations that a large fraction of repositories are personal or inactive and that most forks never produce contributions, which directly undermines naive readings of fork counts as engagement [12]. These studies establish empirically what this paper takes as a design premise: stars measure attention and forks measure intent, and neither should be read as adoption. The contribution here is to push that premise one step further, from “noisy signal” to “adversarial signal,” following the evidence assembled by He et al. [1].
Composite scoring systems
Several operational systems already compute composite scores over open source repositories. The OpenSSF Scorecard runs automated checks against a repository’s security practices, including branch protection, dependency update tooling, and release signing, and aggregates them into a single score [13]; its check set overlaps substantially with the security and release terms proposed here. The OpenSSF Criticality Score aggregates log-scaled activity and dependency signals to estimate how critical a project is to the broader ecosystem [14], and Libraries.io’s SourceRank combines packaging, documentation, and usage signals into a quality rank used for package discovery [15]. In the research community, Munaiah et al. built reaper, a classifier that separates engineered software projects from the long tail of toy and personal repositories using process and documentation signals [16]. The OSCC differs from these systems in three ways. First, it is explicitly a conceptual model rather than a production scorer: the weights are meant to be tuned per evaluation context (DevRel, platform engineering, research) rather than fixed globally. Second, none of the existing scorers includes an explicit suspicion term: Scorecard, Criticality Score, and SourceRank all consume popularity and activity signals at face value, whereas the model proposed here discounts stars and forks by an operationalized manipulation-risk factor. Third, the framework is paired with an evidence hierarchy intended for goal-setting in DevRel, a use case that existing scorers do not address.
Metric manipulation and adversarial measurement
The fourth thread is the most recent. He et al.’s StarScout study provides the first large-scale, longitudinal measurement of suspected fake stars on GitHub, detecting low-activity and lockstep starring behavior across platform metadata from 2019 to 2024, and documenting both a surge of fake-star activity in 2024 and its use beyond short-lived malware repositories, including AI/LLM, blockchain, tooling, and tutorial repositories [1]. GitHub’s own Acceptable Use Policies prohibit inauthentic engagement, rank abuse such as automated starring or following, secondary markets for inauthentic activity, and engagement incentivized by rewards [2], confirming that the platform itself treats these patterns as illegitimate. Practitioner investigations of the surrounding marketplace for stars, followers, and aged accounts provide useful illustrative context [3][4], though they are treated here as narrative background rather than as evidentiary backbone; the replication artifacts of the StarScout study [5] and derived tooling such as RealStars [6] show how these detection ideas can be turned into lightweight review workflows. This empirical work lands on older theoretical ground: Goodhart’s law, in Strathern’s canonical formulation, holds that when a measure becomes a target it ceases to be a good measure [17]. Stars became a target the moment they started influencing discovery, funding, and procurement; the StarScout findings are the predictable consequence. This paper imports that lesson into the design of the model itself, both through the suspicion factors of Section 5 and through the reflexivity discussion of Section 8.
Sustainability, abandonment, and bus factor
Finally, the penalty terms of the model draw on the literature on project sustainability. Avelino et al. proposed practical algorithms for estimating the truck factor (bus factor) of real repositories and showed that a majority of popular systems depend on a very small set of developers [18], motivating the concentration penalty in the OSCC. Coelho and Valente studied why modern open source projects fail and identified maintainer burnout, lack of time, and project obsolescence as dominant causes, with failure frequently preceded by observable decay in maintenance activity [19]. Valiev et al. showed that sustained activity in package ecosystems is shaped by ecosystem-level factors such as dependency position and organizational backing, not only by repository-local signals [20], which motivates the inclusion of an ecosystem term alongside repository-level metrics. The recency-weighted core of the model is a direct formalization of the intuition these studies support: credibility is earned continuously and decays when investment stops.
Stars Are an Adversarial Metric
An adversarial metric is a metric that actors can optimize against, manipulate, or arbitrage once it becomes valuable. GitHub stars influence visibility, social proof, developer confidence, investor attention, and sometimes internal prioritization. That makes them attractive targets.
The StarScout study gives this argument an empirical foundation. The paper presents a global and longitudinal measurement study of suspected fake stars on GitHub, covering platform metadata between 2019 and 2024. Its authors built StarScout to detect anomalous starring behavior, including stars from low-activity accounts and coordinated lockstep starring across repositories. The study reports that fake-star activity surged in 2024 and that fake stars were used not only for short-lived malware repositories but also for AI/LLM, blockchain, tooling, and tutorial repositories [1]. This matters for open source measurement because it changes the status of stars: they are not just incomplete; they are gameable, and they are being gamed at scale.
GitHub’s Acceptable Use Policies prohibit inauthentic interactions, rank abuse such as automated starring or following, secondary markets for inauthentic activity, and engagement incentivized by rewards such as tokens, credits, gifts, or giveaways [2]. The manipulation pattern is therefore not merely analytically suspicious; it conflicts with the platform’s own rules. Around these rules, an informal market nonetheless exists in which stars, followers, promotion services, and aged accounts are packaged and sold as credibility [3][4]. These secondary sources are illustrative rather than evidentiary, but they document the incentive system that any popularity-based metric must survive.
The practical conclusion is the design premise of this paper: stars should be treated as weak evidence until they survive an adversarial credibility check. The same reasoning applies, with lower stakes, to forks. A fork is not usage, adoption, or contribution; it may signal intent to inspect, modify, mirror, or contribute, or it may signal nothing [12]. Forks are nevertheless useful as coherence signals: a repository with a very high star count but almost no watchers, forks, issues, dependents, or contributors deserves more scrutiny than one where stars are accompanied by visible usage and contribution traces. Tools inspired by the StarScout research, such as RealStars, combine fork/star and watcher/star ratios with timing analysis and stargazer-profile quality for exactly this purpose [6]. These heuristics help prioritize review; they do not prove intent.
The Open Source Credibility Coefficient
Design principles and notation
The model is built around five principles. First, no single metric should dominate: every component is normalized to the unit interval before weighting, so the aggregate is a bounded combination of bounded evidence. Second, recent evidence matters more than ancient evidence: time-dependent health signals are aggregated with exponentially decaying recency weights. Third, popularity is admitted but discounted: stars and forks enter the score only after being scaled by an estimated authentic-engagement fraction. Fourth, structural risks subtract: maintainer concentration, abandonment, and unaddressed vulnerabilities enter as bounded penalties rather than as missing bonuses. Fifth, weights are contextual: the model is a parameterized family of scores, and different evaluators (a DevRel team, a platform-engineering team, a research group) are expected to instantiate it differently. This last principle is where the model’s goal-awareness lives. The components naturally separate into quality signals (the health core and the penalties) and adoption signals (, , , and the popularity terms), and the weight vector is the evaluator’s explicit statement of how their definition of success mixes the two: a team chasing virality may legitimately drive and toward zero, while a team selecting a load-bearing dependency will do the opposite. The model does not decide what success means; it forces whoever uses it to say so out loud.
Throughout the paper, quality indicators written are rubric scores in , assigned either by automated checks (in the spirit of OpenSSF Scorecard [13]) or by structured human review. Saturating normalizations of unbounded counts use reference scales, written with a subscript, which encode “how much of this signal counts as a lot” for the evaluation context; the saturating log transform
maps any nonnegative count into while preserving the intuition that the difference between 10 and 1,000 matters more than the difference between 50,000 and 100,000.
Unlike an earlier draft of this model, which integrated unnormalized signals over continuous time inside a sigmoid, the formulation below is discrete, normalized, and explicitly calibrated. The discrete formulation matches what an evaluator would actually compute (metrics per monthly window), the normalization prevents the aggregate from saturating the squashing function, and the explicit time origin removes ambiguity about what is being integrated.
The headline equation
Let be a repository (or, more generally, a project spanning one or more repositories), observed over monthly windows indexed by , where is the most recent complete month and larger denotes months further in the past. The Open Source Credibility Coefficient is
where is the logistic function, and are calibration parameters discussed below, and the raw aggregate is
with normalized recency weights
All component terms () lie in . The health weights satisfy and the block weights satisfy alongside the implicit weight of the health block, so that by construction. Because is bounded, the calibration pair can be chosen so that the logistic function operates in its sensitive range rather than saturating: centers the score (a natural default is the midpoint of the achievable range) and sets how sharply the model separates projects. Without this calibration step, a squashing function over unnormalized signals would pin almost every project to a score near 0 or 100, which is one of the failure modes the earlier draft of this model exhibited.
Table 1 summarizes the notation.
| Symbol | Meaning | Range |
|---|---|---|
| Maintenance quality in month | ||
| Release health in month | ||
| Issue responsiveness in month | ||
| Security maturity in month | ||
| Documentation and governance quality in month | ||
| Real usage and adoption in month | ||
| Project maturity signals (normalized) | ||
| Ecosystem fit signals (normalized) | ||
| Suspicion-discounted star signal | ||
| Suspicion-discounted fork signal | ||
| Maintainer-concentration (bus factor) penalty | ||
| Abandonment penalty | ||
| Vulnerability penalty | ||
| Suspicion factors for stars and forks | ||
| Health-signal weights (sum to 1) | ||
| Block weights | ||
| Penalty weights | ||
| Recency decay rate | ||
| Calibration gain and offset | --- |
Table: Notation summary for the Open Source Credibility Coefficient.
Figure 1 summarizes the same structure visually, in plain language, for readers who want the shape of the model before its details.
{width=88%}
The recency-weighted health core
The first block of says that credibility accumulates over a project’s observable history, but that recent evidence matters more than ancient evidence. The weights decay exponentially with the age of the window: a project that was thriving five years ago but has shown no meaningful activity since should not continue receiving full credit forever. Open source trust has a half-life, and encodes it ( gives a half-life of months). This is consistent with the empirical finding that project failure is typically preceded by observable decay in maintenance activity [19].
The six health signals are weighted by through , which allow organizations to tune the model to their goals. A DevRel team launching a new project might prioritize adoption and responsiveness (); a platform-engineering team consuming the project as a dependency might prioritize maintenance and security (); a research group might add a citation-based usage signal inside . The point is that success depends on context, and the model makes that dependence explicit instead of hiding it inside a fixed global formula.
Maintenance quality
Here is the number of active maintainers in month , is the maintainer count considered fully sufficient for the project class, is the median time to first substantive response on pull requests, and is a tolerance constant. The square root encodes diminishing returns: going from one maintainer to three is a significant improvement; going from fifty to sixty is not. The response-time factor rewards projects that process pull requests in a reasonable amount of time. A project does not need to merge every pull request; it needs to demonstrate that contributors are being heard, because silence is often the strongest signal that a community is struggling.
A deliberate change from the earlier draft: maintainer concentration no longer appears inside . The previous formulation multiplied by , which forced for any solo-maintained project regardless of its actual responsiveness, and double-counted the same risk already captured by the bus-factor penalty. Concentration risk is real, but it is a structural risk rather than a monthly activity signal, so it now lives exclusively in the penalty term (Section 4.5).
Release health
Here counts releases in the trailing twelve months as of month . This term asks whether users can understand what changed, upgrade safely, and trust the artifacts: scores versioning predictability, scores the usefulness of release notes, and scores artifact verifiability. The quality factors are combined as a weighted sum rather than a product. This is intentional: in a multiplicative form, a single missing practice (unsigned releases, say, which describes the majority of real-world projects) annihilates the entire term, which is too strong a claim. A missing practice should attenuate release health, not erase it. The activity factor is saturating because frequent releases are good only up to a point; a project that releases constantly with poor communication may actually increase risk for adopters.
Issue responsiveness
This term does not reward a project for having zero issues, which can equally mean that the project is perfect, that nobody uses it, that users report problems elsewhere, or that the tracker is disabled. Instead it rewards engagement: the ratio captures whether maintainers keep up with incoming issues, and the response-time factor penalizes issues left untouched for long periods.
One refinement matters here. The numerator counts issues resolved, defined as closed with a linked fix, a released change, or a substantive maintainer response, rather than merely closed. A raw closed/opened ratio rewards a pathological behavior that this paper elsewhere identifies as a red flag: closing everything immediately, possibly with a bot. Even the resolved-based definition remains imperfect and partially gameable (a sufficiently motivated actor can automate plausible-looking responses), a limitation taken up in Section 8.
Security maturity
This term asks whether the project is prepared when something goes wrong: is there a clear security policy, does the project respond appropriately to disclosed vulnerabilities, does it publish a software bill of materials, and are its dependencies updated and monitored. The deeper a project sits within a technology stack, the more weight this block deserves. These checks deliberately mirror a subset of the OpenSSF Scorecard check families [13], and an implementation could source them from Scorecard directly.
Documentation and governance
This block measures whether the project is understandable and approachable: a clear README, working examples, API documentation, contribution guidelines, and an unambiguous license. Documentation is often overlooked when measuring success, yet it is one of the stronger predictors of approachability and adoption. Documentation is not decoration; it is part of the product.
Real usage
where counts downstream dependents, counts package downloads (npm, NuGet, PyPI, Maven, crates.io, container pulls, installer downloads), and counts verified real-world integrations (vendor packaging, presence in distributions, citations, documented production deployments). This is where the model measures adoption rather than attention. These signals are imperfect: downloads are inflated by CI systems, dependents may be transitive, and integration counts require curation. Even so, they are generally closer to reality than stars, and they are substantially more expensive to fake.
Maturity and ecosystem terms
The maturity score aggregates slow-moving signals that do not belong in a monthly window: project age, number of stable major versions, governance model, maintainer track record, backward-compatibility discipline, and funding transparency. The ecosystem score aggregates fit signals: compatibility with common tooling, availability through package managers, integrations with major frameworks, community extensions, and presence in distributions or curated catalogs. The inclusion of an ecosystem term reflects the finding that sustained project activity is shaped by ecosystem position and organizational backing, not only by repository-local behavior [20]. A successful project is rarely just a repository; it becomes part of an ecosystem.
Penalties
Three structural risks subtract from the aggregate. All three are bounded in , so a penalty can drag a score down without mathematically annihilating every other form of evidence; how much it drags is an explicit modeling choice () rather than an accident of functional form.
Maintainer concentration. With the estimated number of contributors whose departure would leave the project unable to sustain itself [18],
A project maintained by one exceptional individual can still be fragile. This is not a criticism of solo maintainers; it is a recognition of operational risk, and empirical studies suggest the majority of popular projects carry it [18].
Abandonment. With the number of days since the last meaningful commit and a tolerance constant,
The earlier draft used , which grows without bound and eventually dominates every other term, pinning the final score to zero for any sufficiently old project regardless of its history; the saturating form preserves the intended semantics (staleness erodes trust, with diminishing marginal effect) while keeping the penalty commensurate with the bonuses. The qualifier meaningful matters: a typo fix should not reset the abandonment clock, while bug fixes, feature work, release preparation, security patches, compatibility updates, issue triage, and substantive documentation work should. The tolerance is context-dependent: some software is effectively finished and requires little maintenance, while other categories require continuous investment.
Open vulnerabilities. With a normalized severity and the days since disclosure of vulnerability ,
The penalty grows with severity and with how long a disclosed vulnerability has remained unaddressed. Note the design choice: a vulnerability does not automatically destroy trust; a slow or absent response does. A project that discloses, communicates, and patches promptly accumulates almost no penalty, which is the behavior the model intends to reward.
The suspicion-discounted popularity terms
The most important design choice in the model is how stars and forks enter it:
where are estimated inauthentic fractions of the star and fork counts, derived from the suspicion factors of Section 5. Stars and forks are included, but they are triply weakened: the counts are first multiplied by the estimated authentic fraction, then passed through the saturating log (so the difference between 10 and 1,000 stars matters much more than the difference between 50,000 and 100,000), and finally weighted by and , which the evidence hierarchy of Section 6 argues should be small.
This formulation improves on the earlier draft, which divided the count by : with a suspicion factor bounded in , division discounts the count by at most half, far too weak for repositories where measurement studies suggest the overwhelming majority of stars can be inauthentic [1]. Multiplying by an estimated authentic fraction is both stronger and more interpretable, and it connects the model directly to detection research: StarScout-style tooling estimates precisely this kind of fraction [1][5].
A project does not become credible because people clicked a button. It becomes credible because people use it, contribute to it, and depend on it. For DevRel teams, this distinction is consequential: a campaign that generates 5,000 stars may look impressive, but a campaign that generates 50 active contributors, 20 production deployments, or 10 maintained downstream integrations may create far more long-term value.
Operationalizing the Suspicion Factors
Star suspicion
The earlier draft included a suspicion factor but left it abstract. A more useful version makes it explicit. Define
where each signal lies in , so that and can serve directly as the inauthentic-fraction estimate (or, more conservatively, as an input to a calibrated estimator trained against labeled data such as the StarScout dataset [5]). Table 2 defines the signals.
| Term | Meaning | Typical evidence |
|---|---|---|
| Low-activity account signal | Share of stargazers with little visible GitHub activity | |
| Lockstep behavior signal | Groups of accounts starring the same repositories in close temporal windows | |
| Burst signal | Sudden star spikes without a corresponding release, announcement, vulnerability fix, media event, or community event | |
| Stargazer quality | Account age, contribution history, repositories, followers, profile completeness, non-trivial activity | |
| Cross-metric coherence | Consistency between stars and forks, watchers, issues, pull requests, contributors, dependents, and package downloads | |
| Trend exploitation signal | Suspicious star growth concentrated around discovery surfaces such as trending pages or promotional campaigns |
Table: Components of the star suspicion factor . The first two signals correspond directly to the low-activity and lockstep detectors of the StarScout study [1].
Two caveats govern interpretation. First, open source communities are heterogeneous: educational repositories, curated lists, demos, conference samples, and documentation repositories can exhibit legitimately unusual patterns compared with framework or infrastructure repositories, so the reference behavior against which coherence is judged should be class-conditional. Second, and more fundamentally, is a risk amplifier, not a conviction engine. Its output should be read not as “this repository has fake stars” but as “the star count is not coherent with the rest of the evidence; investigate before using it as a success, trust, or funding signal.” Not every young account is fake, not every burst is manipulation, and not every low fork-to-star ratio proves fraud.
Fork suspicion
The earlier draft announced a fork suspicion factor without defining it. Symmetry is restored as follows:
where is the share of forks with no commits after forking; measures backflow, the extent to which forks produce pull requests, maintained derivatives, or downstream packages; and measures the coherence of the fork count with contributor, issue, and dependency activity. A high dormant share is common and only mildly suspicious on its own (most forks are passive [12]); it is the combination of a large fork count, near-zero backflow, and incoherence with other engagement that raises , and with it .
An Evidence Hierarchy for Open Source Metrics
If the model of Section 4 is compressed into a single practical message, it is that not all open source metrics belong at the same level of evidence. Table 3 proposes a hierarchy.
| Evidence level | Example signals | What they usually mean |
|---|---|---|
| Weak awareness signals | Stars, followers, raw forks, page views, social mentions | Attention, curiosity, visibility, or low-friction approval |
| Observable engagement signals | Issues, discussions, pull requests, watchers, meaningful comments | People are spending time with the project |
| Adoption signals | Package downloads, downstream dependents, container pulls, extension installs, template reuse | The project is being tried, installed, or integrated |
| Durable community signals | Repeat contributors, contributor retention, external maintainers, recurring issue reporters | The project is becoming socially sustainable |
| Production and ecosystem signals | Customer references, production deployments, dependency in critical systems, integrations in other credible projects | The project has moved from interest to dependency |
| Trust and resilience signals | Security policy, signed releases, SBOM, predictable releases, bus factor, governance, license clarity | The project can be depended on responsibly |
Table: An evidence hierarchy for open source metrics, ordered from weakest to strongest. Levels lower in the table are harder to fake, slower to accumulate, and more predictive of long-term value.
Two properties order the hierarchy. Cost of falsification increases as one moves down the table: stars can be purchased by the thousand [1][3], whereas repeat external contributors, production references, and signed-release discipline are expensive to counterfeit at scale. Proximity to value also increases: attention is at best a leading indicator of trial, trial of adoption, adoption of dependency. The OSCC weights should follow the hierarchy: and (the popularity weights) small, (usage) and the structural-trust terms large.
Application: writing DevRel objectives and key results
The hierarchy matters most when writing objectives and key results, because a key result is a metric that has been deliberately turned into a target, which is exactly the condition under which Goodhart’s law activates [17].
A weak key result reads: “Reach 5,000 GitHub stars.” It sits entirely at the weakest evidence level, it is purchasable, and a team measured on it is incentivized to optimize attention rather than value.
A better key result pairs visibility with harder-to-fake evidence: “Reach 5,000 GitHub stars, 25 external contributors, 10 merged external pull requests, 20 downstream dependents, and a median first-response time below 5 business days.” Here stars function as a guardrail-adjacent awareness measure while the substance of the key result lives lower in the hierarchy.
The strongest formulations align the metric with the project’s intended role. For a developer tool: active users, repeat usage, issues carrying production context, and package adoption. For a library: downstream dependents, semantic-versioning discipline, compatibility, and release quality. For a platform sample: successful deployments, template reuse, issue quality, and documentation feedback. For a community project: contributor retention, time to second contribution, discussion depth, and maintainer growth. For a critical infrastructure project: security posture, governance, release signing, dependency hygiene, and bus factor. The point is not to ignore stars; it is to demote them to where they belong, as weak awareness evidence, and to ensure that no key result can be satisfied at the awareness level alone.
Illustrative Application
To illustrate how the framework separates cases that a star count conflates, consider four archetypes drawn from patterns documented in the measurement literature [1][12][19]. The instantiation here is qualitative; a quantitative instantiation against real repositories, using the StarScout dataset [5] for labeled suspicious cases, is left as future work and would be the natural next step in validating the model.
The four archetypes are chosen to cover the quality—adoption plane introduced in Section 1, as Table 4 shows. A star count collapses this plane onto its horizontal axis (and an inflatable proxy of it, at that); the OSCC keeps both dimensions visible and lets the evaluator’s weights decide how they trade off.
| Low apparent adoption | High apparent adoption | |
|---|---|---|
| High quality | D: the excellent solo project | C: the boring load-bearing library |
| Low or immature quality | B: the hyped-then-abandoned project (post-decay) | A: the inflated showcase; also the viral-before-mature project |
Table: The four archetypes positioned on the quality—adoption plane. Note that the same cell can host very different cases: an inflated showcase and a genuinely viral but operationally immature project share a quadrant, and the suspicion machinery of Section 5 is what separates them.
Archetype A: the inflated showcase. A repository with 40,000 stars, a large share of stargazers with empty profiles, star growth in sharp bursts uncorrelated with releases, few watchers, near-zero dependents, and no external contributors. The hierarchy places all of its strength at the weakest level; is high, so strips most of the star count, and with and no durable community signals, the OSCC is low despite the headline number. This is precisely the profile that StarScout-style detection flags [1].
Archetype B: the hyped-then-abandoned project. A genuinely popular project whose stars are authentic but whose meaningful activity stopped eighteen months ago. The suspicion factors stay low, but the recency weights discount its historical health, the abandonment penalty approaches saturation, and unresolved issues depress . Its score decays with time, which matches the intuition that trust has a half-life and the finding that failure follows observable decay [19].
Archetype C: the boring load-bearing library. A project with 800 stars, three maintainers, monthly releases with disciplined changelogs, a security policy, thousands of downstream dependents, and presence in distribution package sets. Almost all of its strength lies in the lower, harder-to-fake levels of the hierarchy: , , , and are high while is modest. The OSCC ranks it far above Archetype A, inverting the star ordering — which is the central claim of the paper in miniature.
Archetype D: the excellent solo project. A responsive, well-documented, well-released project maintained by one person. Under the earlier draft, the concentration term inside maintenance quality forced its health core to zero; under the revised model it scores well on , , , and , and carries a visible but bounded penalty through . The model now says what an honest reviewer would say: this is a good project with a real continuity risk, not a worthless one.
A fifth case is worth distinguishing from Archetype A, because it shares a quadrant while being its opposite in kind: the genuinely viral, operationally immature project. Here the star explosion is authentic — growth correlates with releases and media events, stargazers are real accounts, and stars cohere with rapidly rising usage and contributors — so stays low and the popularity and usage terms score high, while security maturity, release discipline, and the bus factor lag behind. Such a project may be an unambiguous success on its own terms. The OSCC does not deny that; it decomposes it. The score reports strong authentic adoption alongside weak structural trust, which is exactly the information an adopter needs to decide whether to ride the wave or wait a few releases. This is the clearest illustration that the model measures credibility, not success: the two can diverge in both directions, and the component readout matters more than the aggregate number.
The Model Is Itself a Goodhart Target
A paper arguing that stars decayed into an adversarial metric because they became a target must apply the same reasoning to its own proposal. If the OSCC, or any composite like it, were widely adopted for funding, procurement, or platform ranking decisions, every component would become a target in turn: bots can close issues with plausible-looking responses, empty releases can be cut on schedule with templated changelogs, SBOMs can be generated without dependency hygiene behind them, and sock-puppet “external contributors” can submit trivial pull requests. Strathern’s formulation of Goodhart’s law admits no exemption for well-intentioned metrics [17].
Three design choices mitigate, without eliminating, this exposure. First, the model deliberately over-weights signals with high falsification cost: sustained downstream dependency, repeat external contributors, and verified production usage are expensive to counterfeit at scale, unlike stars. Second, the suspicion machinery of Section 5 generalizes: the same coherence logic that audits stars can audit issue-closure patterns or release cadence. Third, and most importantly, the model is positioned as a review framework whose weights vary by evaluator, not as a single public leaderboard; a heterogeneous, partly human-in-the-loop evaluation surface is a much poorer manipulation target than one global number. This is the deeper reason the paper does not propose the OSCC as a universal ranking algorithm: not modesty, but mechanism design.
Limitations and Threats to Validity
Construct validity. The model claims to measure “credibility,” an inherently contested construct. The rubric scores embed evaluator judgment, the reference scales encode what counts as “a lot” of each signal, and the weights encode priorities; two reasonable evaluators can produce different scores for the same repository. This is by design (the model is a parameterized family, not an oracle), but it means OSCC values are comparable only under a fixed instantiation. A deeper limit follows from the construct distinction drawn in Section 1: credibility is not success, and any scalar — this one included — collapses the quality—adoption plane onto a single axis. Projects should therefore be ranked only under a fixed, goal-matched instantiation of the weights, and even then the component readout carries more decision value than the aggregate number. The OSCC cannot adjudicate whether a viral-but-immature project “beat” a polished-but-niche one, because that comparison has no answer independent of the goals against which each is judged.
Data availability. Several signals are not uniformly observable. Clone and traffic statistics are visible only to repository owners, so external evaluators cannot use them. Dependent counts are reliable only in ecosystems with resolvable manifests, and they conflate direct and transitive dependency. Download counts are inflated by CI systems and mirrors. Stargazer-level analysis requires API access at a scale that rate limits can make impractical, and StarScout-style detection is itself an estimate with false positives and false negatives [1]. Any implementation must document which signals were actually available.
Population heterogeneity. Reference behavior differs sharply across repository classes: curated lists, tutorials, and conference demos legitimately accumulate stars without dependents, while internal-infrastructure projects accumulate dependents without stars. Suspicion signals and reference scales should be conditioned on project class, and prior work on separating engineered projects from the long tail [16] suggests this classification is itself nontrivial.
Parameter arbitrariness. The paper provides no fitted values for the weights, decay rate, tolerances, or calibration pair, and makes no claim that a single best instantiation exists. Fitting the parameters against expert judgments or outcome data (project survival, vulnerability response in practice) is open empirical work.
No quantitative validation. Section 7 is illustrative, not evaluative. The framework has not been run against a labeled corpus, and its discriminative value relative to simpler baselines (Scorecard alone, SourceRank alone, or stars alone) is unmeasured. The honest claim of this paper is therefore conceptual: it offers a structured, adversarial-aware way to reason about open source credibility, and a vocabulary for better goal-setting, not a validated predictor.
Author positionality. The author is employed by Microsoft, which owns GitHub. The paper criticizes the most visible GitHub metric and relies on GitHub’s published policies; readers should weigh that proximity, in both directions, when assessing the argument.
Using the Framework in Practice
The OSCC is not intended to automatically accept or reject open source projects; it is intended to structure a review. For a non-critical developer tool, a lightweight instantiation may reduce to five questions: is the project maintained, is the license clear, are releases recent, are issues answered, and is the documentation usable. For a critical dependency, the review goes deeper: who maintains it and what is the bus factor; how have past vulnerabilities been handled; are releases signed; who depends on it downstream; is there commercial backing or foundation governance; could we fork or patch it ourselves if needed; and what is the migration path if the project dies. In both cases the formula’s role is to make sure the conversation covers every block of evidence — health, maturity, ecosystem, risk, and only then popularity — rather than to replace judgment with a number.
Conclusion
Open source success cannot be reduced to stars, especially when stars can be purchased, exchanged, or inflated through coordinated behavior at the scale recent measurement studies document [1]. Stars measure attention. Forks measure intent. Downloads measure usage, imperfectly. And no scalar measures success, because success is goal-relative: a project can go viral ahead of its maturity or quietly fulfill its purpose without ever trending. What an evaluator can measure is credibility, and real credibility emerges from a broader pattern: maintenance, responsiveness, release discipline, documentation, security maturity, ecosystem adoption, and resilience over time — weighted by the goals the evaluation is meant to serve.
This paper contributed an adversarial framing for popularity metrics, a goal-aware and explicitly calibrated credibility model in which popularity enters only after a suspicion discount, an operationalization of that discount for both stars and forks, and an evidence hierarchy designed to keep DevRel objectives anchored in signals that are expensive to fake. It also acknowledged the reflexive limit of the exercise: any score that becomes a target inherits the pathology it was built to escape, which is why the model is offered as a review framework rather than a leaderboard.
So the next time someone says, “this project has 40,000 stars, so it must be safe,” the right answer is: perhaps — but let us look at the rest of the evidence.
\begin{equation*} \mathfrak{C}_{\mathrm{OSS}}(G) ;\neq; \mathrm{stars}(G) \end{equation*}
And that is exactly the point.
Acknowledgments {-}
This paper started as a conversation with colleagues while planning the measurement strategy of an upcoming open source project. The author thanks Yohan Lasorsa, Todd Anglin, and Cedric Vidal for the discussions that shaped the argument and for their feedback on earlier drafts of this work. Any remaining errors are the author’s own.
Conflict of Interest {-}
The author is an employee of Microsoft Corporation, which owns GitHub. This paper criticizes the most visible GitHub popularity metric and relies on GitHub’s publicly documented policies; readers should weigh that proximity, in both directions, when assessing the argument. The analysis is based exclusively on publicly available sources and independent academic research; no internal GitHub data was used, and the views expressed are personal and do not necessarily represent the views of Microsoft.
References {-}
[1] Hao He, Haoqin Yang, Philipp Burckhardt, Alexandros Kapravelos, Bogdan Vasilescu, and Christian Kästner. “Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware.” arXiv:2412.13459, submitted December 18, 2024, revised September 6, 2025. https://arxiv.org/abs/2412.13459
[2] GitHub Docs. “GitHub Acceptable Use Policies.” GitHub Site Policy. https://docs.github.com/en/site-policy/acceptable-use-policies/github-acceptable-use-policies
[3] Elena Marchetti. “Inside GitHub’s Fake Star Economy.” Awesome Agents, April 13, 2026. https://awesomeagents.ai/news/github-fake-stars-investigation/
[4] dev_tips. “The fake GitHub economy no one talks about: Stars, Followers, and $5k Accounts.” DEV Community, September 2025. https://dev.to/dev_tips/the-fake-github-economy-no-one-talks-about-stars-followers-and-5k-accounts-43pn
[5] Hao He. “StarScout: Find suspicious (and possibly faked) GitHub stars at-scale.” GitHub repository. https://github.com/hehao98/StarScout
[6] mercurialsolo. “RealStars: Detect fake GitHub stars using CMU StarScout research.” GitHub repository. https://github.com/mercurialsolo/realstars
[7] Kevin Crowston and James Howison. “Assessing the Health of Open Source Communities.” IEEE Computer, 39(5), 2006.
[8] Slinger Jansen. “Measuring the Health of Open Source Software Ecosystems: Beyond the Scope of Project Health.” Information and Software Technology, 56(11), 2014.
[9] CHAOSS Project, Linux Foundation. “Community Health Analytics in Open Source Software.” https://chaoss.community
[10] Hudson Borges and Marco Tulio Valente. “What’s in a GitHub Star? Understanding Repository Starring Practices in a Social Coding Platform.” Journal of Systems and Software, 146, 2018.
[11] Hudson Borges, Andre Hora, and Marco Tulio Valente. “Understanding the Factors that Impact the Popularity of GitHub Repositories.” Proceedings of the 32nd IEEE International Conference on Software Maintenance and Evolution (ICSME), 2016.
[12] Eirini Kalliamvakou, Georgios Gousios, Kelly Blincoe, Leif Singer, Daniel M. German, and Daniela Damian. “The Promises and Perils of Mining GitHub.” Proceedings of the 11th Working Conference on Mining Software Repositories (MSR), 2014.
[13] Open Source Security Foundation. “OpenSSF Scorecard.” https://github.com/ossf/scorecard
[14] Open Source Security Foundation. “OpenSSF Criticality Score.” https://github.com/ossf/criticality_score
[15] Libraries.io. “SourceRank.” https://libraries.io
[16] Nuthan Munaiah, Steven Kroh, Craig Cabrey, and Meiyappan Nagappan. “Curating GitHub for Engineered Software Projects.” Empirical Software Engineering, 22(6), 2017.
[17] Marilyn Strathern. “‘Improving Ratings’: Audit in the British University System.” European Review, 5(3), 1997.
[18] Guilherme Avelino, Leonardo Passos, Andre Hora, and Marco Tulio Valente. “A Novel Approach for Estimating Truck Factors.” Proceedings of the 24th IEEE International Conference on Program Comprehension (ICPC), 2016.
[19] Jailton Coelho and Marco Tulio Valente. “Why Modern Open Source Projects Fail.” Proceedings of the 11th Joint Meeting on Foundations of Software Engineering (ESEC/FSE), 2017.
[20] Marat Valiev, Bogdan Vasilescu, and James Herbsleb. “Ecosystem-Level Determinants of Sustained Activity in Open-Source Projects: A Case Study of the PyPI Ecosystem.” Proceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), 2018.